PubMed HealthSearch

SEARCH · PubMed Health

Results for “PacBio”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Assembling genomes of non-model plants: A case study with evolutionary insights from Ranunculus (Ranunculaceae).

Whereas genome sequencing and assembly technologies are improving, cost can still be prohibitive for plant species with large, complex genomes. As a consequence, genomics work on some taxa in evolutionarily pivotal positions in the vascular plant tree of life has been hampered. The species-rich genus Ranunculus (Ranunculaceae) is an important angiosperm group for the study of polyploidy, apomixis, and reticulate evolution. However, neither mitochondrial nor high-quality nuclear genome sequences are available. This limits phylogenomic, functional, and taxonomic analyses thus far. Here, we tested Illumina short-read, Oxford Nanopore Technology (ONT) and PacBio (HiFi) long-read, and hybrid-read assembly strategies. We sequenced the diploid progenitor species R. cassubicifolius (R. auricomus species complex) and selected the best assemblies in terms of completeness, contiguity, and quality scores. We first assembled the plastome (156 kbp, 85 genes) and mitogenome (1.18 Mbp, 40 genes) sequences using Illumina and Illumina-PacBio-hybrid strategies, respectively. We also present an updated plastome and the first mitogenome phylogeny of Ranunculaceae, including studies of gene loss (e.g., infA, ycf15, or rps) with evolutionary implications. For the nuclear genome sequence, we favored a PacBio-based assembly polished three times with filtered short reads and subsequently scaffolded into eight pseudochromosomes by chromatin conformation data (Hi-C). We obtained a haploid genome sequence of 2.69 Gbp, with 94.1% complete BUSCO genes found and 35 482 annotated genes, and inferred ancient gene duplications compared to existing Ranunculales genomes. The genomic information presented here will enable advanced evolutionary-functional analyses for the species complex, but also for the genus and beyond Ranunculaceae.

Ranunculus

Chromosome level genome assembly and full-length transcriptome of blacktip trevally (Caranx heberi).

Caranx heberi (Bennett, 1830) commonly known as the blacktip trevally belongs to the family Carangidae and is a potential brackishwater aquaculture species. However, the limited genomic resources are hindering the efforts to study its genetic traits and their molecular basis. To bridge this gap, we generated a high-quality reference genome employing multiple sequencing strategies including PacBio Hifi reads (135x), Illumina short reads (150x), and Hi-C chromosome conformation capturing (180x). The high-quality genome assembly consisted of 159 scaffolds summing to 618.71 Mb and an N50 value of 26.72 Mb. Among these, 24 chromosome level scaffolds covered 97.5% of the total assembly. The genome contained 20.94% of repeat elements and 30,354 protein encoding genes. In addition, full-length transcriptomes were generated using the PacBio IsoSeq approach from seven tissues (gill, kidney, liver, muscle, heart, spleen, and intestine). The comprehensive genomic and transcriptomic resources developed in this study will facilitate the domestication and aquaculture development of C. heberi, as well as support research on its nutritional potential, ecological adaptations, and evolutionary biology.

Animals

De novo clustering of large long-read transcriptome datasets with isONclust3.

MOTIVATION: Long-read sequencing techniques can sequence transcripts from end to end, greatly improving our ability to study the transcription process. Although there are several well-established tools for long-read transcriptome analysis, most are reference-based. This limits the analysis of organisms without high-quality reference genomes and samples or genes with high variability (e.g. cancer samples or some gene families). In such settings, analysis using a reference-free method is favorable. The computational problem of clustering long reads by region of common origin is well-established for reference-free transcriptome analysis pipelines. Such clustering enables large datasets to be split roughly by gene family and, therefore, an independent analysis of each cluster. There exist tools for this. However, none of those tools can efficiently process the large amount of reads that are now generated by long-read sequencing technologies. RESULTS: We present isONclust3, an improved algorithm over isONclust and isONclust2, to cluster massive long-read transcriptome datasets into gene families. Like isONclust, isONclust3 represents each cluster with a set of minimizers. However, unlike other approaches, isONclust3 dynamically updates the cluster representation during clustering by adding high-confidence minimizers from new reads assigned to the cluster and employs an iterative cluster-merging step. We show that isONclust3 yields results with higher or comparable quality to state-of-the-art algorithms but is 10-100 times faster on large datasets. Also, using a 256 Gb computing node, isONclust3 was the only tool that could cluster 37 million PacBio reads, which is a typical throughput of the recent PacBio Revio sequencing machine. AVAILABILITY AND IMPLEMENTATION: https://github.com/aljpetri/isONclust3.

Algorithms

NextLongIso: a comprehensive Nextflow pipeline for multi-dimensional long-read RNA-seq analysis.

SUMMARY: Long-read RNA sequencing technologies, including Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT), enable direct characterization of full-length transcripts and transcriptome complexity. However, analysis of long-read RNA-seq data remains fragmented across multiple tools, limiting the ability to obtain a unified view of transcript structure, expression, and regulatory variation in long-read transcriptomes. We present NextLongIso, a scalable and reproducible Nextflow pipeline that enables coordinated analysis of multiple layers of transcript regulation. Rather than focusing solely on transcript reconstruction, NextLongIso integrates transcript discovery with downstream regulatory analyses to jointly characterize alternative splicing, isoform switching, transcript boundary dynamics (including alternative promoters and polyadenylation), and transposable element-associated transcription from both PacBio and ONT datasets. By eliminating complex cross-tool data harmonization, this unified framework facilitates the transition from transcript identification to functional interpretation of transcriptomic variation. AVAILABILITY AND IMPLEMENTATION: NextLongIso is implemented in Nextflow and is freely available at github: https://github.com/YidanSunResearchLab/nf-LongIso.git and Zenodo: https://doi.org/10.5281/zenodo.21049837.

Software

VicMAG, an open-source tool for visualizing circular metagenome-assembled genomes highlighting bacterial virulence and antimicrobial resistance.

Bacterial pathogens spread in clinical and environmental settings, and mobile genetic elements (MGEs), such as plasmids and phages, mediate the transfer of virulence factor genes (VFGs) and antimicrobial resistance genes (ARGs) among bacterial communities. Metagenomic analysis of environmental and wastewater samples using highly accurate long-read sequencing technologies, such as Pacific Biosciences (PacBio) HiFi sequencing, provides valuable insights into monitoring the regional spread of VFGs and ARGs, including dissemination mediated by MGEs. No visualization tool is currently available for the comprehensive display of numerous resulting circular metagenome-assembled genomes (cMAGs) with functional gene annotations. Here, we developed visualization of circular metagenome-assembled genome (VicMAG), a visualization tool for highly complex cMAGs derived from long-read metagenome assemblies annotated using updated databases of VFGs, ARGs, and MGEs. Using 353 cMAGs from PacBio HiFi sequencing of a wastewater sample, we demonstrated the utility of VicMAG for metagenome visualization. VicMAG provides comprehensive, size-aware visualization of cMAGs representing bacterial chromosomes and plasmids, annotated with VFGs, ARGs, and phages. By simultaneously visualizing all cMAGs in a framework, VicMAG facilitates a holistic understanding of the distribution and genomic context of VFGs and ARGs across complex microbial communities. This tool supports integrated surveillance of bacteria associated with virulence and antimicrobial resistance across clinical, environmental, and One Health contexts.

Metagenome

Patterns of Drug Resistance, Drug Resistance Conferring Mutations and Genomic DNA Methylation Revealed in Mycobacterium tuberculosis From South Africa.

Tuberculosis remains a major public health threat globally, with drug-resistant strains undermining treatment efficacy. We analyzed 126 Mycobacterium tuberculosis (M. tuberculosis) isolates with diverse drug resistance spectra and selected 35 for whole genome sequencing (WGS) using Illumina NextSeq, SMRT PacBio Onso and SMRT PacBio Revio sequencing platforms. The study aimed to characterize drug resistance profiles, compare short- and long-read sequencing performance, identify lineages among South African isolates, detect known drug resistance mutations and their lineage-specific patterns, and utilize long-read SMRT platforms for epigenetic profiling. Multiple drug resistance mutations were identified, some lineage-specific, and notably, East-African-Indian (EAI) Lineage 1 isolates often considered less pathogenic, showed significant potential for multidrug-resistance development, including higher fluoroquinolone resistance as compared to other lineages. Three DNA motifs with methylated adenines, namely CACGCaG, CtCCaG and GaTNNNNRtAC, were detected, with methylation patterns varying by lineage and strain due to mutations in the corresponding methyltransferases (MTases). A particularly notable finding was the stable maintenance of a genetic heterogeneity in the mamB MTase, performing methylation at CACGCaG motifs. These results highlight the combined role of genetic and epigenetic variation in M. tuberculosis adaptive evolution and underscore the value of integrating long-read sequencing into TB surveillance and research.

Mycobacterium tuberculosis

Long-read sequencing reveals putatively mobilizable resistance genes and multi-drug resistance plasmids underestimated by short-read metagenomics.

While shotgun metagenomics is often used to profile antibiotic resistome in gut microbial communities, few studies have investigated if the choice of sequencing platform and assembly strategy affect what mobile genetic elements and antimicrobial resistance genes are recovered. In this study, we compared three platforms (Illumina, Oxford Nanopore, and PacBio HiFi) and seven assembly strategies on gut metagenomes from cattle, pig, and human as case studies. Long-read assemblies recovered 5- to 7-fold more plasmid sequence than Illumina in cattle and pig (mean 17.0 Mb vs. 3.1 Mb), while Illumina performed comparably in the less diverse human gut where high per-species coverage enabled effective short-read plasmid assembly. Long reads also detected more resistance genes on plasmid contigs. Hybrid assembly results depended on the algorithm: scaffolding-based OPERA-MS preserved long-read contiguity and recovered more plasmid-borne resistance genes, while the short-read-centric metaSPAdes hybrid mode produced fragmented assemblies. After collapsing haplotype redundancy, PacBio HiFi identified 2 and 49 unique multi-drug resistance plasmid lineages in cattle and pig, respectively. On the other hand, only 2 and 4 were identified from Illumina. Long reads also placed far more ARGs in a putative mobilization context (50-73%) compared to 14-21% for short reads. Platform and assembly strategy are thus key variables in mobilome and resistome characterization and should be accounted for in antimicrobial resistance surveillance.

Animals

Chromosome-scale assembly with improved annotation provides insights into breed-wide genomic structure and diversity in domestic cats.

INTRODUCTION: Comprehensive genomic resources offer insights into biological features, including traits/disease-related genetic loci. The current reference genome assembly for the domestic cat (Felis catus), Felis_Catus_9.0 (felCat9), derived from sequences of the Abyssinian cat, may inadequately represent the general cat population, limiting the extent of deducible genetic variations. OBJECTIVES: The goal was to develop Anicom American Shorthair 1.0 (AnAms1.0), a reference-grade chromosome-scale cat genome assembly. METHODS: In contrast to prior assemblies relying on Abyssinian cat sequences, AnAms1.0 was constructed from the sequences of more popular American Shorthair breed, which is related to more breeds than the Abyssinian cat. By combining advanced genomics technologies, including PacBio long-read sequencing and Hi-C- and optical mapping data-based sequence scaffolding, we compared AnAms1.0 to existing Felidae genome assemblies (20 scaffolds, scaffolds N50 > 150 Mbp). Homology-based and ab initio gene annotation through Iso-Seq and RNA-Seq was used to identify new coding genes and splice variants. RESULTS: AnAms1.0 demonstrated superior contiguity and accuracy than existing Felidae genome assemblies. Using AnAms1.0, we identified over 1.5 thousand structural variants and 29 million repetitions compared to felCat9. Additionally, we identified > 1,600 novel protein-coding genes. Notably, olfactory receptor structural variants and cardiomyopathy-related variants were identified. CONCLUSION: AnAms1.0 facilitates the discovery of novel genes related to normal and disease phenotypes in domestic cats. The analyzed data are publicly accessible on Cats-I (https://cat.annotation.jp/), which we established as a platform for accumulating and sharing genomic resources to discover novel genetic traits and advance veterinary medicine.

Animals

The complete and annotated mitochondrial genome of Hemileia vastatrix Race I, causal agent of coffee leaf rust.

Hemileia vastatrix is the fungal pathogen responsible for coffee leaf rust (CLR), the most economically important disease of Coffea arabica worldwide. Recently, the nuclear genome of this fungus was completely deciphered. However, the mitochondrial genome of H. vastatrix has remained undercharacterized. Here, we present the complete, circularized mitochondrial genome of H. vastatrix Race I (isolate HvRI), assembled using a hybrid approach combining PacBio HiFi long reads and BGIseq short reads. The genome is 173,525 bp in length with a GC content of 33.1% and encodes 41 functional genes, including 15 protein-coding genes, 2 rRNAs, and 24 tRNAs. The assembly reveals significant structural complexity, driven by intron expansion in the cox1 and cob genes. Notably, the atp8 gene contains a group II intron, rare for this locus, whose internal open reading frame displays evidence of pseudogenization via internal stop codons.. We also characterized a putative replication initiation zone (~1.2 kb) defined by a poly-G homopolymer and conserved regulatory motifs. The mitogenome of the HvRI isolate does not contain cob mutations that lead to amino acid substitutions G143A and F129L associated with the quinone outside inhibitor (QoI) fungicide resistance. This high-quality mitogenome is an important resource for comparative mitogenomics, population diversity studies, and the molecular surveillance of QoI fungicide resistance.

Genome, Mitochondrial

Chromosome-Scale Genome of Zoonotic Eyeworm Thelazia callipaeda from China.

Thelazia callipaeda is a vector-borne zoonotic eyeworm infecting companion animals, wildlife, and humans, but chromosome-scale genomic resources from Chinese clinical material remain limited. We generated a genome supported by Pacific Biosciences (PacBio) high-fidelity (HiFi) sequencing and high-throughput chromosome conformation capture (Hi-C) from 100 adult worms recovered from naturally infected dogs in Beijing and compared its chromosome-scale organization with Portuguese assembly GCA_965194785.1. The final assembly spans 119.53 megabases (Mb) and comprises 115 top-level sequences, including four pseudomolecules totaling 91.26 Mb (76.34%) and 111 unanchored sequences. Genome-mode Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis recovered 98.5% complete chromadorean orthologues, and the representative 11,788-protein gene set recovered 92.6%. Sequence-level alignment resolved Chinese chromosomes 1-4 (chr1-chr4) to Portuguese chr1, chrX, chr3, and chr2, respectively, with retained alignments covering 95.9-99.2% of each Chinese pseudomolecule and estimated sequence identities of 99.75-99.91%. Strong chromosome-scale collinearity was accompanied by localized reverse-collinear regions, including 0.243 Mb and 0.115 Mb intervals on chr2-chrX and chr3-chr3. The anchored sequences contained 96.7% of predicted genes and were substantially more gene-dense than the unanchored sequences. These results establish a clinically sourced Chinese chromosome-scale reference and provide a validated framework for future individual-worm, population-genomic, structural-variation, and comparative genomic studies of this parasite.

Hi-C

T2T Genome Assembly and Multi-Omics Data Reveal Terrestrial Adaptation and Mucus Biosynthesis in Tropical Leatherleaf Slug (Laevicaulis alte).

Laevichaulis alte is a slug in the order Systellommatophora that evolved from aquatic ancestors and now faces strong challenges from desiccation, respiration on land, and novel pathogens. Its mucus is essential for water retention, locomotion, and defense. To link terrestrial adaptation with mucus biosynthesis, we generated a gap-free genome assembly of L. alte using PacBio HiFi reads, Oxford Nanopore ultra-long reads, and Hi-C data. The genome shows low heterozygosity and holocentromeric chromosomes. Functional metabolomics revealed marked metabolic shifts between L. alte and the closely related aquatic species Peronia verruculata. In L. alte, differential metabolites were enriched in lipid metabolism, immune regulation, and stress response pathways, consistent with life in a dry and microbe-rich terrestrial environment. Comparative genomics and transcriptomics identified candidate genes linked to mucus secretion and physiological adaptation, including VEGF, ASGR2, and COL6A6. Further analyses highlighted the vascular endothelial growth factor (VEGF) gene family as a key regulator connecting angiogenesis, tissue remodeling, and mucus production pathways in L. alte. Together, this gap-free genome and multi-omics dataset establish a molecular framework that links genomic innovation, mucus biology, and terrestrial adaptation in Systellommatophora, and they offer a basis for understanding ecological niche specialization in land molluscs.

Animals

Reference-Guided Chromosome-Scale Genome Assembly With Insights on Population Genomics of the Atlantic Goliath Grouper (Epinephelus itajara), Islas del Rosario, Colombia.

Epinephelus itajara, commonly known as the Atlantic Goliath grouper, is the largest species among the western North Atlantic groupers and is critically endangered. This species plays a crucial ecological, cultural, and economic role and has been the focus of captive breeding efforts at the Oceanario of the Rosario Islands, Colombia. However, despite its ecological and conservation importance, genomic resources and population genomic data for E. itajara remain scarce, particularly in the Colombian Caribbean. This study presents a reference-guided chromosome-scale genome assembly and an analysis of the population genomic structure of E. itajara using PacBio HiFi sequencing and Illumina technologies. The assembled genome has a total size of 1.12 Gb, with a contig N50 of 42.69 Mb and a scaffold N50 of 46.30 Mb. A total of 22,692 protein-coding genes were identified after masking 46% of the genome, which consists of repetitive elements. Comparative genomic analyses revealed a high degree of collinearity with closely related Epinephelus species and identified E. lanceolatus as the closest relative, supporting recent divergence and conserved genome architecture within the genus. Additionally, a population genomics analysis was conducted using 7706 high-quality SNPs to assess the genomic structure of captive populations. The results revealed four distinct genomic lineages, with moderate genetic differentiation among the sampled individuals. In the Colombian Caribbean, two unique lineages were identified, associated with the localities of Bahía Cispatá and Bahía Barbacoas, suggesting possible geographic isolation. These genomic resources provide valuable tools and new opportunities to better understand the genomic diversity, evolutionary history, and reproductive mechanisms of E. itajara. Moreover, they serve as a foundation for conservation strategies, including selective breeding programs aimed at increasing genomic diversity in captive populations and guiding restoration efforts in its natural habitat.

Epinephelus itajara

First clinical diagnosis of FAME3 via commercial Long-Read sequencing reveals mosaic repeat expansion in MARCHF6 gene.

Familial Adult Myoclonic Epilepsy type 3 (FAME3) is a rare autosomal dominant disorder characterized by cortical tremor and epilepsy, caused by a noncoding pentanucleotide repeat expansion (TTTTA/TTTCA)n in the MARCHF6 gene. Conventional genetic testing often fails to detect this expansion due to its repetitive structure and intronic location. We evaluated a 61-year-old woman with refractory myoclonic and generalized tonic-clonic seizures, whose prior genetic testing-including exome and genome sequencing-was non-diagnostic. Using PacBio HiFi long-read whole-genome sequencing and the tandem repeat genotyping tool TRGT, we identified a pathogenic MARCHF6 intronic expansion. The proband harbored one allele with 15 TTTTA repeats and a second allele with a compound expansion of 661 TTTTA and 12 TTTCA repeats. Three affected relatives shared similarly expanded alleles, but with increasing repeat size in the latter generations. Importantly, analysis using TRGT-instability revealed repeat mosaicism in all affected individuals, reflected by variability in motif counts across individual sequencing reads. This somatic heterogeneity may contribute to the phenotypic penetrance, variable expressivity and pleiotropism seen in FAME3 disease expression. To our knowledge, this is the first clinical diagnosis of FAME3 using a commercially available long-read sequencing platform, underscoring its diagnostic utility in resolving complex repeat expansion disorders and uncovering biologically relevant mosaicism.

Humans

Insights into dill (Anethum graveolens) flavor formation via integrative analysis of chromosomal-scale genome, metabolome and transcriptome.

INTRODUCTION: Dill (Anethum graveolens) is a significant medicinal herb belonging to the Apiaceae family. Owing to its high levels of volatile organic compounds (VOCs), dill is commonly utilized for essential oil extraction and medicine purpose. However, the biosynthesis of the crucial VOC in dill remains obscure. OBJECTIVES: Identify the key VOCs related to the flavor formation in dill and dissect the regulatory mechanism of their synthesis. METHODS: The dill chromosomal-level genome was constructed by PacBio HiFi, Hi-C, and BGISEQ second generation sequencing and assembly. The VOCs in dill leaves were identified through GC-MS. The potential mechanism involved in regulating the VOC accumulation in dill flavor formation was analyzed by multi-omics analysis. RESULTS: A 1.17 Gb chromosome-scale genome of dill with a contig N50 of 10.78 Mb was constructed. A total of 46,538 genes were annotated across 11 assembled chromosomes. Comparative genomics analysis suggested that transposable element insertions, especially LTR-Gypsy, have contributed to the evolution and expansion of the dill genome. The flavor formation of dill was mainly attributed to terpenoids, especially α-phellandrene, β-ocimene, and o-cymene. The contribution of expansion and replication of terpenoid synthesis pathway genes, especially terpene synthase (TPS), to the abundant terpenoid production of dill was identified. Differential gene expression patterns observed at various developmental stages and tissues provided key candidate genes for the regulation of terpenoid synthesis, as well as transcription factors. The different accumulation of esters and aromatics also affected the flavor formation of dill. The key genes implicated in the synthesis of anethole, namely AIS and AMT were further identified. CONCLUSION: This study constructed the chromosome level genome and identified the main VOCs and related key genes in flavor formation of dill, shedding lights on our understanding of terpenoid biosynthesis but also offered guidance for future genetic research on molecular breeding in Anethum graveolens.

Transcriptome

Salmonella Pullorum strain SPullorum-YN-07 from dead embryos of Yanjin black-bone chickens: Complete genome with IncFII(S) and Col(pVC) plasmids and pathogenicity.

Salmonella Pullorum is a host-adapted pathogen that causes Pullorum disease in chickens and can be vertically transmitted via eggs, leading to embryonic mortality. The susceptibility and vertical transmission of S. Pullorum may vary among chicken breeds, yet genomic characterization of strains from dead embryos of indigenous breeds remains limited. This study isolated and characterized a Gram-negative short rod, designated Salmonella Pullorum strain SPullorum-YN-07, from dead embryos of Yanjin black-bone chickens, a native breed in Yunnan, China. The strain formed colorless colonies on MacConkey agar and red, non-H2S colonies on XLD agar, with biochemical reactions consistent with the genus Salmonella. Whole-genome sequencing using Illumina and PacBio platforms generated a complete genome consisting of one circular chromosome and four circular plasmids; plasmid replicon types IncFII(S) and Col(pVC) were identified in two of the plasmids. On the chromosome, a total of 340 virulence-associated genes were detected, including those involved in secretion systems, adhesion, motility, and immune modulation. Resistance gene analysis identified the acquired aminoglycoside resistance gene aac(6')-Iaa, alongside multiple intrinsic resistance determinants related to efflux pumps and target alteration. Multilocus sequence typing (MLST) assigned the strain to sequence type ST92, and core-genome phylogenetic analysis confirmed its clustering within the Salmonella Pullorum lineage. In a chick infection model, the strain induced depression, white diarrhea, and growth retardation, with clinical scores peaking at 10 days post-infection and a mortality rate of 10%. Bacterial colonization was highest in the cecum, and histopathological lesions were observed in the liver, spleen, and cecum. This study provides the first complete genomic characterization and pathogenicity assessment of an S. Pullorum strain isolated from dead embryos of Yanjin black-bone chickens, offering a foundation for understanding host-pathogen interactions in indigenous breeds and assessing cross-transmission risks to commercial poultry populations.

Complete genome

Genome assembly of Astatotilapia latifasciata uncovers B chromosome-linked chromatin reorganization.

B chromosomes (Bs) are supernumerary genomic elements found in many eukaryotes, yet their full sequence composition, functional potential, and regulatory impact on the host genome remain unclear. Here, we present a chromosome-level genome assembly of the cichlid fish Astatotilapia latifasciata, integrating PacBio long reads, Illumina short reads, and Hi-C chromatin contact maps to resolve both A and B chromosomes. The 0.93 Gb assembly (N50 = 36.2 Mb) includes a 34 Mb B chromosome containing 789 predicted protein-coding genes and a markedly higher density of transposable elements (TEs), especially long terminal repeats (LTR) retrotransposons. Transcriptome profiling revealed that B-linked genes are predominantly transcriptionally repressed relative to their A chromosome paralogs. Hi-C-based chromatin modeling uncovered distinct 3D structural configurations associated with the B chromosome, including fewer topologically associating domains (TADs), reduced loop formation, and altered compartmentalization. These changes are linked to long-range chromatin interactions and genomic rearrangements, suggesting that the B chromosome reshapes the nuclear architecture of the host genome. Our study proposes a potential regulatory role of Bs in genome and provides a genomic resource for investigating chromosome evolution in cichlids.

Animals

Accurate somatic small variant discovery for multiple sequencing technologies with DeepSomatic.

Somatic variant detection is an integral part of cancer genomics analysis. While most methods have focused on short-read sequencing, long-read technologies offer potential advantages in repeat mapping and variant phasing. We present DeepSomatic, a deep-learning method for detecting somatic small nucleotide variations and insertions and deletions from both short-read and long-read data. The method has modes for whole-genome and whole-exome sequencing and can run on tumor-normal, tumor-only and formalin-fixed paraffin-embedded samples. To train DeepSomatic and help address the dearth of publicly available training and benchmarking data for somatic variant detection, we generated and make openly available the Cancer Standards Long-read Evaluation (CASTLE) dataset of six matched tumor-normal cell line pairs whole-genome sequenced with Illumina, PacBio HiFi and Oxford Nanopore Technologies, along with benchmark variant sets. Across samples, both cell line and patient-derived, and across short-read and long-read sequencing technologies, DeepSomatic consistently outperforms existing callers.

Humans

Chromosome-level genome assembly of Triplophysa scleroptera.

Triplophysa scleroptera is an endemic fish species in Qinghai Lake and the upper reaches of the Yellow River. However, studies on conservation and evolutionary genetics were seriously impeded by the absence of a reference genome. Here, by using PacBio HiFi sequencing and Hi-C assembly technology, we assembled a chromosome-level genome of T. scleroptera, with a total length of 660.22 Mb and 99.82% of the sequence anchored to 25 chromosomes. The contig N50 and scaffold N50 were 9.09 Mb and 24.38 Mb, respectively. The evaluation using BUSCO indicated the genome assembly to be 96.40% complete. About 33.41% of the genome consists of repeat elements. We predicted 26,168 protein-coding genes in the genome, and 99.02% of them were functionally annotated. This high-quality reference genome would serve as a valuable genomic resource for advancing evolutionary conservation genetics studies in this species.

Animals