PubMed HealthSearch

SEARCH · PubMed Health

Results for “reference genome”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Long-read Sequences Mapped to a Complete Reference Genome Uncover Uncaptured Structural Variants across the Beta-globin Cluster in Africans with Sickle Cell Disease.

African genomes are marked by extensive complexity in the number and distribution of variants, yet remain under-represented in genetic databases and the human reference genome. This gap in representation limits the broad application of genomic medicine. Sickle cell disease (SCD) - one of the most common monogenic diseases - has its highest prevalence in Africa, and variation in disease severity has consistently been linked to the beta-globin locus, including levels of fetal hemoglobin (HbF). Modulation of HbF is central to current SCD gene therapies; however, the inherent complexity and variation at the locus in African genomes presents a challenge to translating these advances to Africa. Here, we align long-read single molecule sequences (LRS) targeted to the beta-globin region to the hg38 and T2T-CHM13v2 genome references in 40 individuals with SCD, predominantly recruited from three African countries. We demonstrate that the expanded T2T-CHM13v2 reference sequence at this locus reduces Structural Variant (SV) calls by 70% and uncovers uncaptured single nucleotide variants (SNVs). Across the cluster we report 343 SVs and 196 SNVs that have not been previously reported, including in LRS data from the All of Us project. By including African populations from ethnolinguistic groups that have not been previously surveyed we improve variant resolution and bolster evidence for observed variation. Finally, we identify a common ∼4kb insertion locus overlapping the HBB promoter among individuals with high HbF. These results demonstrate the utility of combining a comprehensive reference genome with LRS in African populations to uncover genomic variation at disease-associated loci.

SNV

JG2: an updated version of the Japanese population-specific reference genome.

Here we present the construction of JG2, an updated population-specific reference genome for the Japanese population. Utilizing data from three individuals previously used in the construction of JG1, several methodologies were employed to enhance genomic coverage and assembly quality. Hi-C sequencing technology facilitated phase-aware assembly, generating two haploid assemblies per individual and enabling improved representation of genetic variation. A meta-assembly strategy and a majority decision approach further refined assembly quality by combining the best sequences from multiple assemblies and minimizing the inclusion of rare variants. The resulting JG2 genome comprises chromosome-level sequences, mitochondrial chromosomes and unplaced scaffolds, offering more comprehensive coverage of the Japanese genome. Comparative analyses with other reference genomes demonstrated the accuracy and representativeness of JG2, highlighting its utility for genetic research involving the Japanese population. Overall, by adopting the phased assembly technique, JG2 represents a substantial advancement over the collapsed assembly-based JG1, with improvements including a greater number of identified variants (3,115,695 variants, of which 298,644 had an allele frequency (AF) of 1.0 in the 3.5KJPNv2 AF panel) and a higher N50 value (152,668,378 bp). These enhancements provide researchers with a more precise and comprehensive resource for understanding the genetic landscape of the Japanese population. The sequences and annotations are available on the jMorp website ( https://jmorp.megabank.tohoku.ac.jp/ ).

Journal Article

ERGA-BGE reference genome of Eunicella cavolini, an IUCN Near Threatened Gorgonian of the Mediterranean Sea.

The Eunicella cavolini reference genome provides an important resource to study the adaptation of this species to different environments and anthropic pressures. This species is impacted by human activities, including climate change, and this reference genome will be useful to study the genomic evolution of this species. The entirety of the genome sequence was assembled into 17 contiguous chromosomal pseudomolecules. This chromosome-level assembly encompasses 0.49 Gb, composed of 159 contigs and 46 scaffolds, with contig and scaffold N50 values of 7.7 Mb and 51.1 Mb, respectively.

Biodiversity Genomics Europe

Reference genome bias in light of species-specific chromosomal reorganization and translocations.

BACKGROUND: Whole-genome sequencing efforts, have during the past decade, unveiled the central role of genomic rearrangements-such as chromosomal inversions-in evolutionary processes, including local adaptation in a wide range of taxa. However, employment of reference genomes from distantly or even closely related species for mapping and the subsequent variant calling can lead to errors and/or biases in the datasets generated for downstream analyses. RESULTS: Here, we capitalize on the recently generated chromosome-anchored genome assemblies for Arctic cod (Arctogadus glacialis), polar cod (Boreogadus saida), and Atlantic cod (Gadus morhua) to evaluate the extent and consequences of reference bias on population sequencing datasets (approx. 15-20 × coverage) for both Arctic cod and polar cod. Our findings demonstrate that the choice of reference genome impacts the mapping statistics, including mapping depth and mapping quality, as well as core population genetic estimates, such as heterozygosity levels, nucleotide diversity (π), and cross-species genetic divergence (DXY). Furthermore, using a more distantly related reference genome can lead to inaccurate detection and characterization of chromosomal inversions, i.e., in terms of size (length) and location (position), due to inter-chromosomal reorganizations between species. Additionally, we observe that some of the verified species-specific inversions are split across multiple genomic regions when mapped against a heterospecific reference. CONCLUSIONS: Inaccurate identification of chromosomal rearrangements as well as biased population genetic measures could potentially lead to erroneous interpretation of species-specific genomic diversity, impede the resolution of local adaptation, and thus, impact predictions of their genomic potential to respond to climatic and other environmental perturbations.

Animals

Reference genome of the Californian trapdoor spider Aptostichus stephencolberti Bond 2008 (Araneae: Mygalomorphae: Euctenizidae).

We present a reference genome assembly for the trapdoor spider Aptostichus stephencolberti. This species, described in 2008, is endemic to the highly fragmented coastal dune habitats of Northern California from Monterey to the San Francisco Bay Area. Trapdoor spiders are ideal taxa for landscape scale genomic studies owing to their extreme site fidelity and limited dispersal capabilities; these same characteristics make them prone to extinction. Genomic studies of species like A. stephencolberti can reveal novel areas of endemism and high conservation value that may not be evident in species with wider ranges and greater dispersal capabilities. As part of the California Conservation Genomics Project, we constructed the A. stephencolberti reference genome from high quality long-read sequences, scaffolded with proximity ligation Omni-C data. The primary assembly comprises 551 scaffolds spanning 3.63 Gbp, a scaffold N50 of 62.2 Mbp and BUSCO completeness of 95.6%. We estimate 52 chromosomes yet find no (TTAGG)n telomer repeats. Expanding the telomeric repeat search finds an ancestral loss of the repeat from all spiders. Automated annotation using the NCBI refseq pipeline and RNAseq data from whole adults finds 14,067 genes with a BUSCO annotation completeness of 95.56%. Repeat annotation identified 77% of the genome to be interspersed repeats. This resource, the first for family Euctenizidae will facilitate future study and resulting conservation actions of A. stephencolberti and other Aptostichus sp. populations associated with the rapidly changing California coastal dune ecosystem.

Aptostichus stephencolberti

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 × homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering ∼18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals

Phased telomere-to-telomere reference genome and pangenome reveal an expansion of resistance genes during apple domestication.

The cultivated apple (Malus domestica Borkh.) is a cross-pollinated perennial fruit tree of great economic importance. Earlier versions of apple reference genomes were unphased, fragmented, and lacked comprehensive insights into the apple's highly heterozygous genome, which impeded advances in genetic studies and breeding programs. In this study, we assembled a haplotype-resolved telomere-to-telomere (T2T) reference genome for the diploid apple cultivar Golden Delicious. Subsequently, we constructed a pangenome based on 12 assemblies from wild and cultivated species to investigate the dynamic changes of functional genes. Our results revealed the gene gain and loss events during apple domestication. Compared with cultivated species, more gene families in wild species were significantly enriched in oxidative phosphorylation, pentose metabolic process, responses to salt, and abscisic acid biosynthesis process. Our analyses also demonstrated a higher prevalence of different types of resistance gene analogs (RGAs) in cultivars than their wild relatives, partially attributed to segmental and tandem duplication events in certain RGAs classes. Structural variations, mainly deletions and insertions, have affected the presence and absence of TIR-NB-ARC-LRR, NB-ARC-LRR, and CC-NB-ARC-LRR genes. Additionally, hybridization/introgression from wild species has also contributed to the expansion of resistance genes in domesticated apples. Our haplotype-resolved T2T genome and pangenome provide important resources for genetic studies of apples, emphasizing the need to study the evolutionary mechanisms of resistance genes in apple breeding.

Malus

ERGA-BGE reference genome of the Eurasian Woodcock ( Scolopax rusticola), a game bird species with isolated populations of conservation interest.

The reference genome of the Eurasian Woodcock ( Scolopax rusticola) is an important resource to investigate population structure across the wide breeding range of this iconic game species and the conservation status of specific management units, such as the isolated Macaronesian populations. The genome sequence was assembled into 45 contiguous chromosomal pseudomolecules and 2 sex chromosomes (W and Z). This chromosome-level assembly encompasses 1.2 Gb, composed of 1,613 contigs and 935 scaffolds, with contig and scaffold N50 values of 5.9 Mb and 34.2 Mb, respectively.

Aves

Near-complete reference genome assembly of Hoya carnosa.

Hoya R. Br. is the largest genus in the tribe Marsdenieae (Apocynaceae), comprising 350-450 species. Hoya species are popular in horticulture for their distinctive floral traits and fragrances, primarily sourced from domestication and mutation breeding. However, the lack of molecular analysis for floral morphological traits has limited their cultivation and application. In this study, we assembled a near-complete reference genome for H. carnosa, the model species of the genus, using PacBio HiFi reads and Hi-C method. The genome size was approximately 465.7 Mb with a contig N50 of 39.3 Mb. 99.7% of the sequences were anchored to 11 pseudochromosomes, and the assembly achieved a BUSCO score of 98.5%. We predicted 24,309 protein-coding genes, of which 90.2% (21,927) were functionally annotated. This high-quality genome provides a valuable reference for the research of evolution, conservation and molecular breeding in Hoya.

Genome, Plant

A complete and near-perfect rhesus macaque reference genome: lessons from subtelomeric repeats and sequencing bias.

A truly complete, telomere-to-telomere (T2T), and error-free reference genome remains a foundational resource-and long-standing goal-for unbiased comparative and functional genomics. While recent T2T assemblies of humans and other primates have made substantial progress, most still contain thousands of base-level errors, particularly within highly repetitive regions. Here, we present T2T-MMU8v2.0, a near-perfect T2T assembly of the rhesus macaque (Macaca mulatta), representing the highest base-level accuracy reported in a primate genome to date. By employing an optimized ONT-only assembly strategy, we identify subtelomeric satellite-rich regions as the principal bottleneck to improving assembly quality, owing to technological biases in long-read platforms and limitations in current hybrid assembly frameworks. We discover 268 previously unannotated repeat families and resolve ~8 Mbp of SATR satellite arrays, with over 99-fold enrichment in historically misassembled subtelomeric regions. These satellites form four distinct genomic architectures, each with unique SATR satellite composition, segmental duplication organization, and epigenetic signatures, distinct from the subtelomeric architectures observed in hominid genomes. Notably, in contrast to the largely gene-poor subtelomeric regions in African hominids, the SATR architectures in macaques harbor 58 actively transcribed genes, supported by open chromatin and expression data, suggesting gene innovation within these repetitive regions. Functionally, T2T-MMU8v2.0 improves read mappability and accuracy across sequencing platforms, and results in a 19% improvement of transcription start site enrichment scores and 5,821 additional chromatin accessibility peaks on average, thereby enhancing variant detection, regulatory annotation, and transcriptomic resolution in population genetics or single-nucleus studies. Together, this work establishes a new benchmark for genomics, offers a roadmap for resolving complex repetitive regions, and reveals previously unrecognized features of subtelomeric genome structure and evolution.

Journal Article

Chromosome-Level Reference Genome of the Desert Night Lizard Xantusia vigilis.

We present a reference-quality genome assembly for the desert night lizard (Xantusia vigilis). The night lizards (Xantusiidae) are a family of small-bodied lizards found in North America (Xantusia), Central America (Lepidophyma), and Cuba (Cricosaura). The night lizard family has an independent evolutionary history of at least 80 million years from its sister taxa within Scincoidea. The Xantusiids have several unique ecological, behavioral and evolutionary characteristics. For instance, the family contains the only squamate species that form diploid, unisexual, parthenogenic lineages. In addition, most night lizards are viviparous and form stable kin groups that are maintained over multiple years, an unusual life history strategy among lizards. Combining PacBio long-read sequencing, Hi-C, and RNAseq data we developed a reference-quality genome for the desert night lizard, X. vigilis. We assembled a complete mitochondrion and ~ 2.2 Gb nuclear genome, with 20 scaffolds that correlate in size to the X. vigilis karyotype. In addition, we found that X. vigilis chromosome 1 aligns with gene content of both of macrochromosome 1 and microchromosome 9 from a genome assembly of a species in the sister family Cordylidae (Hemicordylus capensis).

Xantusia

Reference genomes of Japanese raccoon dog (Nyctereutes viverrinus) and a Japanese red fox (Vulpes vulpes japonica).

We established primary fibroblast cultures from a Japanese raccoon dog (Nyctereutes viverrinus) and a Japanese red fox (Vulpes vulpes japonica) and generated highly contiguous reference genome assemblies using Oxford Nanopore Technologies PromethION long-read sequencing. The Japanese raccoon dog assembly spanned 2.69 Gb in 813 scaffolds, with a scaffold N50 of 52 Mb and a Benchmarking Universal Single-Copy Orthologs (BUSCO) completeness score of 98.2%. The Japanese red fox assembly spanned 2.47 Gb in 903 scaffolds, with a scaffold N50 of 139 Mb and a BUSCO completeness score of 97.5%. Phylogenomic analysis placed the Japanese raccoon dog in a lineage distinct from the continental raccoon dog, supporting its evolutionary differentiation within Nyctereutes. The Japanese red fox formed a distinct lineage within the red fox clade, consistent with its recognized regional differentiation. These genome assemblies and associated fibroblast cultures provide resources for studies of canid systematics, population history, local adaptation, comparative genome evolution, and conservation genetics.

Canidae

ERGA-BGE reference genome of the Mediterranean monk seal ( Monachus monachus), an IUCN Vulnerable species.

The Mediterranean monk seal, Monachus monachus, is the only pinniped that lives in the Mediterranean Sea and one of the rarest marine mammals in the world. The species was recently classified as "vulnerable" by the IUCN, considering an improvement in its overall status. However, the species' populations have undergone severe bottlenecks due to systematic persecution by humans over the past centuries. Today, the global population of M. monachus is estimated to be no more than 1,000 individuals. The Mediterranean monk seal is mainly using marine caves as resting and pupping sites. It is an opportunistic apex predator and as a result its role is considered important for maintaining the structure and function of marine ecosystems. Nowadays, the species is threatened mainly by the destruction of its habitat (due to coastal development, mass tourism, and pollution) and by the depletion of its prey due to overfishing. The Mediterranean monk seal is an emblematic species; its ecological importance and its vulnerable status render its protection and effective management necessary. The entirety of the genome sequence of a female specimen was assembled into 16 contiguous chromosomal pseudomolecules, one sex chromosome (X), and one mitochondrial genome. This chromosome-level assembly encompasses 2.4 Gb, composed of 316 contigs and 275 scaffolds, with contig and scaffold N50 values of 90.9 Mb and 157 Mb, respectively.

Biodiversity Genomics Europe

The reference genome of the human diploid cell line RPE-1.

Recent technological advances have facilitated the assembly of telomere-to-telomere (T2T) genomes. The current T2T CHM13 showcases the complete architecture of the human genome, yet its use in functional experiments is limited by discrepancies with the actual genome of the specific biological system under study. Access to reference assemblies for experimentally relevant cell lines is therefore essential in advancing sequencing-based analyses and precise manipulation, particularly in highly variable regions such as centromeres. Here, we present RPE1v1.1, the near-complete diploid genome assembly of the hTERT RPE-1 cell line, a non-cancerous human retinal epithelial model with a stable karyotype. Using high-coverage Pacific Biosciences and Oxford Nanopore Technologies long-read sequencing, we generate a high-quality de novo assembly, validate it through multiple methods, and phase it by integrating high-throughput chromosome conformation capture (Hi-C) data. Our assembly includes chromosome-level scaffolds that span centromeres for all chromosomes. Comparing both haplotypes with the CHM13 genome, we detect haplotype-specific genomic variations, including the translocation between chromosome 10 and chromosome X t(X;10)(Xq28;10q21.2) characteristic of RPE-1 cells, and divergence peaking at centromeres. Altogether, the RPE1v1.1 genome provides a reference-quality diploid assembly of a widely used cell line, supporting high-precision genetic and epigenetic studies in this model system.

Humans

A chromosome-level reference genome assembly of the Small snakehead (Channa asiatica).

The Small snakehead (Channa asiatica) is an economically important species in both aquaculture and ornamental trade, mainly distributed in South China and Southeast Asia. Despite its significance, limited genomic resources have impeded in-depth genetic studies and breeding programs. In this study, we used PacBio HiFi long-read sequencing, Illumina short-read sequencing, and Hi-C technologies to generate a high-quality chromosome-level genome of the C. asiatica. The final genome spans 659.44 Mb, with an impressive 98.18% anchored to 23 chromosomes. Notably, the contig N50 and scaffold N50 are 23.92 Mb and 29.61 Mb, validated by a BUSCO completeness score of 98.93%. Genome annotation identified 26,603 protein-coding genes, 99.29% of which were confirmed by BUSCO analysis, and 93.68% were functionally annotated. Approximately 27.72% of the genome sequences were classified as repeat elements. This high-fidelity genome assembly provides a robust foundation for advancing molecular breeding, comparative genomics, and evolutionary studies of C. asiatica and related species.

Animals

A telomere-to-telomere reference genome assembly of the red silk cotton tree (Bombax ceiba).

Bombax ceiba, an important ornamental tree and potential fiber resource in the textile industry, is widely distributed in tropical and subtropical regions. In this study, we assembled a nearly gap-free telomere-to-telomere (T2T) genome of B. ceiba using Illumina, PacBio High-fidelity (HiFi), ONT ultra-long, and Hi-C sequencing technologies. The genome spanned approximately 807.89 Mb, with a scaffold N50 of 16.58 Mb, and 754.68 Mb (93.41%) of genomic sequences were anchored onto 48 pseudo-chromosomes. Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis revealed a completeness of 99.40%, identifying 1,378 single-copy and 213 duplicated genes out of 1,614. The genome contained 67.72% (547.11 Mb) repeat regions, with 39,708 predicted protein-coding genes. Collectively, our study provides valuable genomic data for investigating the evolutionary history of the Malvaceae family.

Genome, Plant

Defining three dimensional chromatin structures of pediatric and adolescent B cells using primary B cell and EBV-immortalized B cell reference genomes.

BACKGROUND/PURPOSE: Knowledge of the 3D genome is essential to elucidate genetic mechanisms driving autoimmune diseases. The 3D genome is distinct for each cell type, and it is uncertain whether cell lines faithfully recapitulate the 3D architecture of primary human cells or whether developmental aspects of the pediatric immune system require use of pediatric samples. We undertook a systematic analysis of B cells and B cell lines to compare 3D genomic features encompassing risk loci for juvenile idiopathic arthritis (JIA), systemic lupus (SLE), and type 1 diabetes (T1D). METHODS: We isolated B cells from four healthy individuals, ages 9-17. HiChIP was performed using a CTCF antibody, and CTCF peaks were called within each sample separately. Peaks observed in all four samples were identified. CTCF loops were called within the pediatric samples using three CTCF peak datasets: 1) self-called CTCF consensus peaks called within the pediatric samples, 2) ENCODE's publicly available GM12878 CTCF ChIP-seq peaks, and 3) ENCODE's primary B cell CTCF ChIP-seq peaks from two adult females. Differential looping was assessed within the pediatric samples and each of the three peak datasets. RESULTS: The number of consensus peaks called in the pediatric samples was similar to that identified in ENCODE's GM12878 and primary B cell datasets. We observed&#x2009;<&#x2009;1% of loops that demonstrated significantly differential looping between peaks called within the pediatric samples themselves and when called using ENCODE GM12878 peaks. Significant looping differences were even fewer when comparing loops of the pediatric called peaks to those of the ENCODE primary B cell peaks. When querying loops found in juvenile idiopathic arthritis, type 1 diabetes, or systemic lupus erythematosus risk haplotypes, we observed significant differences in only 2.2%, 1.0%, and 1.3% loops, respectively, when comparing peaks called within the pediatric samples and ENCODE GM12878 dataset. The differences were even less apparent when comparing loops called with the pediatric vs ENCODE adult primary B cell peak datasets. CONCLUSION: The 3D chromatin architecture in B cells is similar across pediatric, adult, and EBV-transformed cell lines. This conservation of 3D structure includes regions encompassing autoimmune risk haplotypes. Thus, even for pediatric autoimmune diseases, publicly available adult B cell and cell line datasets may be sufficient for assessing effects exerted in the 3D genomic space.

Humans