PubMed HealthSearch

SEARCH · PubMed Health

Results for “PacBio sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Salmonella Pullorum strain SPullorum-YN-07 from dead embryos of Yanjin black-bone chickens: Complete genome with IncFII(S) and Col(pVC) plasmids and pathogenicity.

Salmonella Pullorum is a host-adapted pathogen that causes Pullorum disease in chickens and can be vertically transmitted via eggs, leading to embryonic mortality. The susceptibility and vertical transmission of S. Pullorum may vary among chicken breeds, yet genomic characterization of strains from dead embryos of indigenous breeds remains limited. This study isolated and characterized a Gram-negative short rod, designated Salmonella Pullorum strain SPullorum-YN-07, from dead embryos of Yanjin black-bone chickens, a native breed in Yunnan, China. The strain formed colorless colonies on MacConkey agar and red, non-H2S colonies on XLD agar, with biochemical reactions consistent with the genus Salmonella. Whole-genome sequencing using Illumina and PacBio platforms generated a complete genome consisting of one circular chromosome and four circular plasmids; plasmid replicon types IncFII(S) and Col(pVC) were identified in two of the plasmids. On the chromosome, a total of 340 virulence-associated genes were detected, including those involved in secretion systems, adhesion, motility, and immune modulation. Resistance gene analysis identified the acquired aminoglycoside resistance gene aac(6')-Iaa, alongside multiple intrinsic resistance determinants related to efflux pumps and target alteration. Multilocus sequence typing (MLST) assigned the strain to sequence type ST92, and core-genome phylogenetic analysis confirmed its clustering within the Salmonella Pullorum lineage. In a chick infection model, the strain induced depression, white diarrhea, and growth retardation, with clinical scores peaking at 10 days post-infection and a mortality rate of 10%. Bacterial colonization was highest in the cecum, and histopathological lesions were observed in the liver, spleen, and cecum. This study provides the first complete genomic characterization and pathogenicity assessment of an S. Pullorum strain isolated from dead embryos of Yanjin black-bone chickens, offering a foundation for understanding host-pathogen interactions in indigenous breeds and assessing cross-transmission risks to commercial poultry populations.

Complete genome

The chromosome-level genome assembly and annotation of the silver-lipped pearl oyster, Pinctada maxima.

The silver-lipped pearl oyster (Pinctada maxima) is a valuable tropical aquaculture species, playing a crucial economic role in the global pearl industry. However, the lack of genomic reference limits our in-depth understanding of this species in genome-based breeding, conservation, evolution and adaptation. Here, annotated chromosome-level reference genome for P. maxima was generated by integrating PacBio long-read sequencing, Illumina short-read sequencing, and Hi-C sequencing data. The total genome size is 1,264.93&#x2009;Mb, with contig N50 and scaffold N50 of 649&#x2009;kb and 89.19&#x2009;Mb, respectively. The majority (97.94%) of the assembled genome was anchored to the 14 chromosomes by Hi-C analysis. The relatively high genome completeness was observed, with 97.38% (metazoa_odb10 database) and 95.26% (mollusca_odb10 database) in BUSCO analysis. Genome annotation revealed approximately 65.46% of the repeat sequences and 26,315 protein-coding genes. Comparative genome analysis revealed 28 expanded and 48 contracted families (p&#x2009;<&#x2009;0.05) in P. maxima, with 3.2% of genes (894) being species-specific. This chromosome-level genome serves as an essential resource for research in evolutionary genomics, phylogenetics, and biomineralization.

Animals

Chromosome-level genome assembly of bivalve mollusk, Xishishe Coelomactra antiquata.

Coelomactra antiquata, a significant marine economic shellfish in China, is experiencing a natural population decline due to habitat destruction and overfishing, making the restoration and conservation of its natural resources an urgent priority. This study provides a high - quality chromosome - level genome assembly for C. antiquata, created by PacBio and Hi - C sequencing and resulting in a 19 - chromosome map. The assembly encompasses a genome size of 807.31&#x2009;Mb, with a contig N50 of 17.35&#x2009;Mb and a scaffold N50 of 42.90&#x2009;Mb. A total of 28,070 protein - coding genes were identified, 25,959 of which were functionally annotated. Overall, this study offers a chromosome - level genome for C. antiquata that is highly continuous and complete, providing an indispensable resource for subsequent molecular and genetic studies of this species.

Animals

Chromosome-Level Reference Genome of the Desert Night Lizard Xantusia vigilis.

We present a reference-quality genome assembly for the desert night lizard (Xantusia vigilis). The night lizards (Xantusiidae) are a family of small-bodied lizards found in North America (Xantusia), Central America (Lepidophyma), and Cuba (Cricosaura). The night lizard family has an independent evolutionary history of at least 80 million years from its sister taxa within Scincoidea. The Xantusiids have several unique ecological, behavioral and evolutionary characteristics. For instance, the family contains the only squamate species that form diploid, unisexual, parthenogenic lineages. In addition, most night lizards are viviparous and form stable kin groups that are maintained over multiple years, an unusual life history strategy among lizards. Combining PacBio long-read sequencing, Hi-C, and RNAseq data we developed a reference-quality genome for the desert night lizard, X. vigilis. We assembled a complete mitochondrion and&#x2009;~&#x2009;2.2 Gb nuclear genome, with 20 scaffolds that correlate in size to the X. vigilis karyotype. In addition, we found that X. vigilis chromosome 1 aligns with gene content of both of macrochromosome 1 and microchromosome 9 from a genome assembly of a species in the sister family Cordylidae (Hemicordylus capensis).

Xantusia

Genetic Adaptation to Brackish Water and Spawning Season in European Cisco.

How species adapt to diverse environmental conditions is essential for understanding evolution and the maintenance of biodiversity. The European cisco (Coregonus albula) is a salmonid that occurs in both fresh and brackish water, and this together with the presence of sympatric spring- and autumn-spawning lacustrine populations provides an opportunity for studying the genetics of adaptation in relation to salinity and timing of reproduction. Here, we present a high-quality reference genome of the European cisco based on PacBio HiFi long read sequencing and HiC-directed scaffolding. We generated low-coverage whole-genome sequencing data from 336 individuals across 12 population samples to explore population structure and genetics of ecological adaptation. We found a major subdivision between two groups of populations most likely reflecting colonisation from different glacial refugia. Within the two major groups, we detected further genetic differentiation between spring- and autumn-spawning populations and between populations from freshwater lakes, rivers and brackish water (Bothnian Bay). A genome-wide screen for genetic differentiation among populations identified a set of outlier SNPs strongly correlated with spawning timing and salinity. Several of the genes associated with spawning time, including BHLHE40, TIMELESS and CPT1A, have previously been shown to have a role in circadian rhythm biology. As many as 17 loci were associated with genetic differentiation between populations reproducing in fresh and brackish water. This study provides insights into the genomic basis of ecological adaptation in European cisco with implications for sustainable fishery management.

Animals

Integrative Long-Read Multi-Omics of a Patient With GPI Deficiency: A Molecular Case Study of a Candidate Dual-Effect GPI Variant.

The molecular determinants of phenotypic severity in red cell enzymopathies are often obscured by the disconnect between coding sequence variants and their regulatory landscapes. Here we present a single-patient molecular case study that uses an integrative multi-omic approach-combining short-read WGS, PacBio HiFi long-read sequencing, native CpG methylation profiling, and Iso-Seq full-length transcriptomics-to characterize a severe, transfusion-dependent hemolytic anaemia. We identified a compound heterozygous state in the glucose-6-phosphate isomerase (GPI) gene, with no wild-type allele present. One allele (Haplotype 1) carried a missense variant (p.His191Arg); the other (Haplotype 2) carried a distinct missense variant, c.1414C>T (p.Arg472Cys), previously reported as biochemically unstable. Long-read phasing placed the two variants in trans. Allele-resolved transcript counts showed a directionally consistent but statistically non-significant trend toward higher expression of Haplotype 2 across two Iso-Seq replicates. Notably, the c.1414C>T transition abolishes a local CpG dinucleotide; in a small number of haplotype-2 reads spanning this position, the corresponding cytosine on the wild-type/Haplotype-1 background was methylated. We did not measure GPI protein abundance, enzymatic activity, or stability in this patient, and we do not establish that methylation at this site regulates GPI transcription. On the basis of these correlative observations in a single patient, we propose-as a hypothesis for future testing-that a coding variant might simultaneously perturb protein stability and disrupt a local epigenetic mark, and we outline the experiments required to test whether such a dual effect contributes to disease. This case illustrates the value of integrative long-read multi-omics for generating mechanistic hypotheses about variants of uncertain significance, while underscoring that causal claims require dedicated functional validation.

Humans

Assembling genomes of non-model plants: A case study with evolutionary insights from Ranunculus (Ranunculaceae).

Whereas genome sequencing and assembly technologies are improving, cost can still be prohibitive for plant species with large, complex genomes. As a consequence, genomics work on some taxa in evolutionarily pivotal positions in the vascular plant tree of life has been hampered. The species-rich genus Ranunculus (Ranunculaceae) is an important angiosperm group for the study of polyploidy, apomixis, and reticulate evolution. However, neither mitochondrial nor high-quality nuclear genome sequences are available. This limits phylogenomic, functional, and taxonomic analyses thus far. Here, we tested Illumina short-read, Oxford Nanopore Technology (ONT) and PacBio (HiFi) long-read, and hybrid-read assembly strategies. We sequenced the diploid progenitor species R. cassubicifolius (R.&#x2009;auricomus species complex) and selected the best assemblies in terms of completeness, contiguity, and quality scores. We first assembled the plastome (156&#x2009;kbp, 85 genes) and mitogenome (1.18&#x2009;Mbp, 40 genes) sequences using Illumina and Illumina-PacBio-hybrid strategies, respectively. We also present an updated plastome and the first mitogenome phylogeny of Ranunculaceae, including studies of gene loss (e.g., infA, ycf15, or rps) with evolutionary implications. For the nuclear genome sequence, we favored a PacBio-based assembly polished three times with filtered short reads and subsequently scaffolded into eight pseudochromosomes by chromatin conformation data (Hi-C). We obtained a haploid genome sequence of 2.69&#x2009;Gbp, with 94.1% complete BUSCO genes found and 35&#x2009;482 annotated genes, and inferred ancient gene duplications compared to existing Ranunculales genomes. The genomic information presented here will enable advanced evolutionary-functional analyses for the species complex, but also for the genus and beyond Ranunculaceae.

Ranunculus

CLN3 transcript complexity revealed by long-read RNA sequencing analysis.

BACKGROUND: Batten disease is a group of rare inherited neurodegenerative diseases. Juvenile CLN3 disease is the most prevalent type, and the most common pathogenic variant shared by most patients is the "1-kb" deletion which removes two internal coding exons (7 and 8) in CLN3. Previously, we identified two transcripts in patient fibroblasts homozygous for the 1-kb deletion: the 'major' and 'minor' transcripts. To understand the full variety of disease transcripts and their role in disease pathogenesis, it is necessary to first investigate CLN3 transcription in "healthy" samples without juvenile CLN3 disease. METHODS: We leveraged PacBio long-read RNA sequencing datasets from ENCODE to investigate the full range of CLN3 transcripts across various tissues and cell types in human control samples. Then we sought to validate their existence using data from different sources. RESULTS: We found that a readthrough gene affects the quantification and annotation of CLN3. After taking this into account, we detected over 100 novel CLN3 transcripts, with no dominantly expressed CLN3 transcript. The most abundant transcript has median usage of 42.9%. Surprisingly, the known disease-associated 'major' transcripts are detected. Together, they have median usage of 1.5% across 22 samples. Furthermore, we identified 48 CLN3 ORFs, of which 26 are novel. The predominant ORF that encodes the canonical CLN3 protein isoform has median usage of 66.7%, meaning around one-third of CLN3 transcripts encode protein isoforms with different stretches of amino acids. The same ORFs could be found with alternative UTRs. Moreover, we were able to validate the translational potential of certain transcripts using public mass spectrometry data. CONCLUSION: Overall, these findings provide valuable insights into the complexity of CLN3 transcription, highlighting the importance of studying both canonical and non-canonical CLN3 protein isoforms as well as the regulatory role of UTRs to fully comprehend the regulation and function(s) of CLN3. This knowledge is essential for investigating the impact of the 1-kb deletion and rare pathogenic variants on CLN3 transcription and disease pathogenesis.

Humans

Chromosome-Scale Genome of Zoonotic Eyeworm Thelazia callipaeda from China.

Thelazia callipaeda is a vector-borne zoonotic eyeworm infecting companion animals, wildlife, and humans, but chromosome-scale genomic resources from Chinese clinical material remain limited. We generated a genome supported by Pacific Biosciences (PacBio) high-fidelity (HiFi) sequencing and high-throughput chromosome conformation capture (Hi-C) from 100 adult worms recovered from naturally infected dogs in Beijing and compared its chromosome-scale organization with Portuguese assembly GCA_965194785.1. The final assembly spans 119.53 megabases (Mb) and comprises 115 top-level sequences, including four pseudomolecules totaling 91.26 Mb (76.34%) and 111 unanchored sequences. Genome-mode Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis recovered 98.5% complete chromadorean orthologues, and the representative 11,788-protein gene set recovered 92.6%. Sequence-level alignment resolved Chinese chromosomes 1-4 (chr1-chr4) to Portuguese chr1, chrX, chr3, and chr2, respectively, with retained alignments covering 95.9-99.2% of each Chinese pseudomolecule and estimated sequence identities of 99.75-99.91%. Strong chromosome-scale collinearity was accompanied by localized reverse-collinear regions, including 0.243 Mb and 0.115 Mb intervals on chr2-chrX and chr3-chr3. The anchored sequences contained 96.7% of predicted genes and were substantially more gene-dense than the unanchored sequences. These results establish a clinically sourced Chinese chromosome-scale reference and provide a validated framework for future individual-worm, population-genomic, structural-variation, and comparative genomic studies of this parasite.

Hi-C

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software

NextLongIso: a comprehensive Nextflow pipeline for multi-dimensional long-read RNA-seq analysis.

SUMMARY: Long-read RNA sequencing technologies, including Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT), enable direct characterization of full-length transcripts and transcriptome complexity. However, analysis of long-read RNA-seq data remains fragmented across multiple tools, limiting the ability to obtain a unified view of transcript structure, expression, and regulatory variation in long-read transcriptomes. We present NextLongIso, a scalable and reproducible Nextflow pipeline that enables coordinated analysis of multiple layers of transcript regulation. Rather than focusing solely on transcript reconstruction, NextLongIso integrates transcript discovery with downstream regulatory analyses to jointly characterize alternative splicing, isoform switching, transcript boundary dynamics (including alternative promoters and polyadenylation), and transposable element-associated transcription from both PacBio and ONT datasets. By eliminating complex cross-tool data harmonization, this unified framework facilitates the transition from transcript identification to functional interpretation of transcriptomic variation. AVAILABILITY AND IMPLEMENTATION: NextLongIso is implemented in Nextflow and is freely available at github: https://github.com/YidanSunResearchLab/nf-LongIso.git and Zenodo: https://doi.org/10.5281/zenodo.21049837.

Software

Population-scale detection of methylation outliers from long-read genome sequencing.

BACKGROUND: Aberrant DNA methylation can mediate the functional effects of rare genetic variation and contribute to imprinting disorders, repeat expansion diseases, and other pathogenic regulatory mechanisms. Long-read sequencing technologies now enable genome-wide detection of CpG methylation alongside genetic variation from a single assay. However, methods for systematic identification and interpretation of methylation outliers from long-read sequencing data remain limited. METHODS: We developed METAFORA, a computational workflow for detecting methylation outlier regions from PacBio and Oxford Nanopore long-read sequencing data. METAFORA constructs population-level methylation references, segments the genome into correlated CpG blocks, infers technical and biological sources of variation through hidden factor estimation, models uncertainty due to variable depth sequencing, and computes covariate-adjusted methylation outlier scores for individual samples. We applied METAFORA across large long-read sequencing cohorts and integrated methylation outliers with multi-omic data. METAFORA is implemented as a snakemake workflow available at https://github.com/tjense25/METAFORA. RESULTS: METAFORA identified methylation outlier regions associated with rare structural variants, tandem repeat expansions, and imprinting abnormalities. We found outlier regions were enriched for molecular outliers across transcriptomic and chromatin accessibility datasets, supporting their functional relevance in gene regulation. In a representative case, METAFORA identified an imprinting defect affecting the GNAS locus associated with an STX16 deletion. CONCLUSIONS: METAFORA enables scalable detection and interpretation of methylation outliers from long-read sequencing data and provides a framework for integrating epigenetic outliers with genomic and multi-omic analyses. These approaches may improve interpretation of rare regulatory variation and support discovery of clinically relevant epigenetic abnormalities in genomic medicine.

DNA methylation

Chromosome-level genome assembly of Cheilinus chlorourus (Bloch, 1791) (Perciformes: Labridae).

In the classification of marine fish, the Labridae family ranks second in terms of species diversity and plays a vital role in coral reef ecosystems, comprising over 600 species across 82 genera. Despite its significance for ecological and evolutionary studies, genomic research on this group has lagged, resulting in a shortage of data, particularly regarding high-quality chromosome-level genome assemblies. To address this gap, this study focused on Cheilinus chlorourus from the Labridae family and successfully achieved a chromosome-level genome assembly. By integrating Illumina, PacBio, and Hi-C sequencing data, we assembled a genome measuring 940.36&#x2009;Mb, with 926.86&#x2009;Mb (98.56%) of the gene assembly organized into 21 chromosomes. A total of 29,213 protein-coding genes (PCGs) were identified, and 79.93% of these genes were functionally annotated. With this high-quality genome assembly, future investigations into the functional genomics and ecology of C. chlorourus will have a solid scientific foundation.

Animals

A chromosome-level reference genome assembly of the Small snakehead (Channa asiatica).

The Small snakehead (Channa asiatica) is an economically important species in both aquaculture and ornamental trade, mainly distributed in South China and Southeast Asia. Despite its significance, limited genomic resources have impeded in-depth genetic studies and breeding programs. In this study, we used PacBio HiFi long-read sequencing, Illumina short-read sequencing, and Hi-C technologies to generate a high-quality chromosome-level genome of the C. asiatica. The final genome spans 659.44&#x2009;Mb, with an impressive 98.18% anchored to 23 chromosomes. Notably, the contig N50 and scaffold N50 are 23.92&#x2009;Mb and 29.61&#x2009;Mb, validated by a BUSCO completeness score of 98.93%. Genome annotation identified 26,603 protein-coding genes, 99.29% of which were confirmed by BUSCO analysis, and 93.68% were functionally annotated. Approximately 27.72% of the genome sequences were classified as repeat elements. This high-fidelity genome assembly provides a robust foundation for advancing molecular breeding, comparative genomics, and evolutionary studies of C. asiatica and related species.

Animals

Chromosome-level genome assembly of the large carpenter bee Xylocopa dejeanii Lepeletier, 1841 (Hymenoptera: Apidae).

Xylocopinae, a diverse bee subfamily comprising over 1,000 bee species, and also a major model system for studying the pollination and evolution of sociality. The lack of chromosome-level genome assembly resources for the Xylocopinae limits our research of their biology and evolution. Here, we provided the first pseudo-chromosomes genome assembly of the Xylocopa dejeanii combined PacBio CLR long reads, Illumina sequences, and Hi-C data. The final genome is 194.44&#x2009;Mb located in 16 chromosomes. Our assembly includes 141 scaffolds, with a scaffold N50 length of 13.15&#x2009;Mb. BUSCO analysis revealed 99.00% completeness. Genome annotation identified 28.27&#x2009;Mb of repetitive elements, 10,970 protein-coding genes, and 432 ncRNAs. This high-quality X. dejeanii assembly advances our understanding of Xylocopinae genomics and provides new insights into bee evolution.

Animals

Chromosome-level genome assembly of hawthorn spider mite, Amphitetranychus viennensis (Acari: Tetranychidae).

The hawthorn spider mite, Amphitetranychus viennensis, is a major pest of orchards and ornamentals in the Palaearctic region, with adaptability and acaricide resistance. The lack of high-quality genomic resources limits understanding of its detoxification mechanisms and the development of RNAi-based pest control strategies. In this study, we utilized Illumina, Pacific Biosciences (PacBio), and Hi-C sequencing technologies to assemble a chromosome-level reference genome of A. viennensis. The assembled genome spans 141.96&#x2009;Mb, with a contig N50 of 1.35&#x2009;Mb. BUSCO analysis confirmed a high level of completeness, covering 91.6% of annotated genes. The assembly includes 50.97&#x2009;Mb of repetitive sequences, representing 35.93% of the genome, and annotates 13,968 protein-coding genes. Using Hi-C sequencing, we anchored 47 contigs to three chromosomes, accounting for 97.27% of the estimated nuclear genome and achieving a contig N50 of 45.83&#x2009;Mb. This high-quality genome assembly provides a valuable foundation for evolutionary and genomic research on spider mites, while also serving as a genetic resource to inform molecular control strategies and support sustainable pest management.

Animals

Chromosomal-level genome assembly of Trypanosoma carassii, the etiologic agent of a recent outbreak of trypanosomiasis in cage-cultured large yellow croaker (Larimichthys crocea) in China.

Trypanosoma carassii, a typical freshwater fish trypanosome, has recently been identified as the etiological agent of a trypanosomiasis outbreak in cage-cultured large yellow croaker (Larimichthys crocea) in China and has been designated as T. c. larimichthys. To date, publicly available genomic data for trypanosomes have been limited to terrestrial species, particularly those of medical importance. Here, we present a chromosome-level genome assembly of T. carassii, the first genome of an aquatic trypanosome, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding technologies. A preliminary genome survey based on Illumina sequencing data estimated the genome size at 56.38&#x2009;Mb with a heterozygosity of 1.17%. The final assembled genome spans 48.55&#x2009;Mb, with contig N50 and scaffold N50 values of 139.15 Kb, and achieves 100.00% BUSCO completeness. Hi-C data resolved the assembly into 34 chromosomes and 9 unanchored scaffolds. Repetitive elements account for 53.29% of the genome (approximately 25.87&#x2009;Mb). A total of 11,584 protein-coding genes were predicted, 95.36% of which were functionally annotated. Synonymous substitution rates analysis of paralogous genes indicates a recent burst of gene duplication, which likely corresponds to a whole-genome duplications. This high-quality genome assembly provides invaluable resources for understanding the evolution and host adaptation of aquatic trypanosomes.

Animals

High-Resolution Chromosome-Level Genome Assembly and Annotation of Triplophysa stewarti, an Endemic Plateau Loach from the Qinghai-Tibet Plateau.

The bottom-dwelling fish Triplophysa stewarti, endemic to the Qinghai-Tibet Plateau, is a valuable model for studying high-altitude adaptation in aquatic ecosystems. However, the lack of a high-quality reference genome has hindered comparative genomic and evolutionary studies within this genus. Here, we present a chromosome-level genome assembly for T. stewarti, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding. The 697.9&#x2009;Mb assembly is highly continuous (scaffold N50 of 253.58&#x2009;Mb) and encompasses 25 chromosomes, representing 92.65% of the genome. BUSCO analysis indicated a 98.4% completeness, supporting the high quality of the assembly. We annotated 28,009 protein-coding genes, with 97.04% being functionally assigned across multiple databases (NR, UniProt, KEGG, GO, Pfam and InterPro). Additionally, repetitive elements constituted 42.47% of the genome, and we identified 52,709 non-coding RNAs. This high-quality reference genome provides a fundamental resource for exploring the adaptive evolution, population structure, and conservation genetics of T. stewarti and related species on the Qinghai-Tibet Plateau.

Animals