PubMed HealthSearch

SEARCH · PubMed Health

Results for “reduced representation sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Enzymatic depletion of transposable elements in sequencing libraries and its application for genotyping multiplexed CRISPR-edited plants.

Whole-genome sequencing has become a common strategy to genotype individual plants of interest. Although a limited number of genomic regions usually need to be surveyed with this strategy, excess sequencing information is almost always generated at an appreciable financial cost. Repetitive sequences (e.g., transposons), which can account for more than 80% of the genome of some plants, are often not required in these genotyping projects. Therefore, strategies that enrich DNA coding for the protein-coding genes prior to sequencing can lower the cost to obtain sufficient sequence information. Here, we present the development and application of methylation-sensitive reduced representation sequencing (MsRR-Seq), which relies on the cytosine methylation-sensitive restriction enzyme MspJI to deplete constitutive heterochromatic DNA before library construction. By applying MsRR-Seq to citrus and maize, we show that protein-coding genes can be enriched in sequencing datasets. We then describe the application of MsRR-Seq to facilitate the identification of complex mutants from populations of citrus plants resulting from multiplex CRISPR/Cas9 editing of four genes. Overall, this work demonstrates an easy and low-cost method to enrich non-repetitive DNA in high-throughput sequencing libraries, an approach that is especially useful for large plant genomes with an excessively high proportion of methylated repetitive sequences.

DNA Transposable Elements

Scalable medium-density genotyping platforms for cultivar identification, pedigree authentication, marker-assisted and genomic selection, and other applications in strawberry.

A broad spectrum of high-density genotyping approaches, including single-nucleotide polymorphism (SNP) arrays, genotyping-by-sequencing, and whole-genome reduced-representation sequencing, have been shown to perform well in strawberry (Fragaria × ananassa), despite the inherent complexity of the octoploid genome. While these approaches are effective, their routine deployment in breeding programs can be constrained by cost, computational requirements, and workflow complexity. In parallel, many breeding programs continue to rely on locus-specific assays for marker-assisted selection, resulting in fragmented and inefficient genotyping strategies. Here, we describe medium-density amplicon-based genotyping platforms for strawberry designed to provide cost-effective, turnkey solutions that integrate markers used for marker-assisted selection with genome-wide markers suitable for genomic prediction in a single laboratory assay. These platforms were developed by targeting 1,650 or 4,811 target SNPs via amplicon sequencing, and are interoperable with existing high-density genotyping resources, including a widely used 50K SNP array, thereby facilitating data integration across platforms. We benchmarked their performance relative to the 50K SNP array across breeding-relevant applications, including identity and purity testing, pedigree authentication, marker-assisted selection, and genomic selection, and further evaluated the feasibility of genotype imputation to enhance genome-wide information content. Across analyses, the 1,650- and 4,811-amplicon platforms produced results comparable to higher-density platforms while substantially reducing genotyping cost and analytical overhead. This work demonstrates that targeted amplicon-based genotyping can support efficient, scalable, and integrated genome-informed breeding, enabling the routine application of both marker-assisted and genomic selection within strawberry breeding workflows. Open-source R workflows are provided to support streamlined analyses in breeding contexts.

Fragaria

Integrative analysis of transcriptome and DNA methylome dynamics during caudal fin regeneration in silver pomfret (Pampus argenteus).

Caudal fin regeneration in teleost fish is a complex, multi-stage process involving coordinated molecular and cellular changes. While the role of epigenetic regulation particularly DNA methylation has been studied in model freshwater species such as zebrafish, its contribution to regeneration in marine teleosts remains largely unexplored. In this study, we integrated transcriptomic and DNA methylomic data to characterize the temporal dynamics of gene expression and methylation during caudal fin regeneration in the silver pomfret (Pampus argenteus). Using RNA-sequencing and reduced representation bisulfite sequencing (RRBS) at three biologically critical time points 1, 3, and 7 days post-amputation (dpa), we characterized the spatiotemporal molecular landscape of caudal fin regeneration. These time points capture the key transitional phases of wound healing and inflammation (1 dpa), blastema formation and progenitor proliferation (3 dpa), and regenerative outgrowth with tissue remodeling (7 dpa), enabling robust detection of the major molecular programs underlying epimorphic regeneration. Concurrently, CG-methylome analysis identified thousands of dynamically changing differentially methylated regions (DMRs). A strong global inverse correlation was observed between promoter methylation and gene expression. Integrative analysis pinpointed key regeneration genes (fgf20a, msxb, sox9b) whose expression was associated with dynamic methylation changes in their promoters or gene bodies. We conclude that DNA methylation is a dynamic and key regulatory layer that acts in concert with transcriptional reprogramming to coordinate tissue regeneration, providing new insights into the epigenetic mechanisms underlying complex regenerative processes in teleosts.

Animals

Climate Gradients and Habitat Discontinuity Structure Genetic Variation in a Spring-Specialist Plant.

BACKGROUND AND AIMS: Groundwater-dependent ecosystems support disproportionate biodiversity in arid regions, yet the population genetics of spring-specialist plants remains poorly understood. Here, we present the first species-wide genetic dataset for crimson monkeyflower (Mimulus verbenaceus, Phrymaceae), a spring-specialist plant distributed in seeps, springs, and associated riparian areas across desert regions of North America. We aim to relate landscape features and climate gradients to the spatial genetic structuring within this system. METHODS: Using genome-wide reduced representation sequencing data consisting of 10,760 SNPs from 175 individuals across 17 populations, we characterized the patterns of genetic differentiation and diversity. Population structure was assessed using ADMIXTURE and Principal Component Analysis. We examined the contributions of climate to range-wide genetic variation in crimson monkeyflower using redundancy analysis. KEY RESULTS: Patterns of genetic differentiation were more consistent with those of spring-specialist animal taxa than those of upland plants or generalist riparian plants. We found strong population structure at both broad regional scales and at fine local scales. While geographic and spatial structuring was a primary driver of genetic structure across all scales, riparian connectivity influenced local patterns of diversity, and adaptation to local climatic variation was more influential at regional scales, with temperature, relative humidity, and a monsoon-driven climate gradient contributing to genetic differentiation. CONCLUSIONS: Our findings highlight the distinctive association with isolated perennial groundwater sources, as well as climate gradients, with genetic variation in this spring-specialist plant. These findings suggest that spring-specialist plants deserve special consideration in ecological theory, management, and conservation.

Mimulus

Epigenetic Profiling for Early Detection and Treatment Response Monitoring in Non-Small Cell Lung Cancer: Protocol for a Prospective Translational Biomarker Study.

BACKGROUND: Non-small cell lung cancer (NSCLC) is the leading cause of cancer-related mortality worldwide and continues to have poor survival outcomes, with most patients diagnosed at advanced stages of disease. In New Zealand, NSCLC contributes substantially to cancer inequities, with Māori communities experiencing disproportionately high incidence and mortality rates. Although low-dose computed tomography screening can improve early detection, major limitations remain, including false-positive findings, overdiagnosis, high infrastructure costs, and limited accessibility for rural and underserved populations. Liquid biopsy approaches using circulating tumor DNA (ctDNA), particularly DNA methylation profiling, have emerged as promising, minimally invasive strategies for improving cancer detection, treatment monitoring, and precision oncology. OBJECTIVE: This study aims to establish integrated genomic and epigenomic predictive and prognostic biomarkers using ctDNA, tumor tissue, and transcriptomic profiling to improve early detection, risk stratification, treatment selection and response prediction, and longitudinal monitoring, with particular emphasis on identifying molecular mechanisms associated with treatment resistance and disease progression. METHODS: This prospective observational translational biomarker study is being conducted through the University of Otago and associated respiratory and oncology services in New Zealand. The study will recruit participants with NSCLC (including squamous and nonsquamous subtypes), individuals referred to fast-track lung nodule assessment clinics, and nonmalignant respiratory controls. Serial peripheral blood sampling will be performed in selected participants at predefined clinical follow-up time points to evaluate treatment response and disease progression. The availability of formalin-fixed paraffin-embedded archival tissues will be recorded, but will not be mandatory for enrollment. Genome-scale DNA methylation profiling will be performed using cell-free reduced representation bisulfite sequencing (cfRRBS), while targeted genomic profiling and transcriptomic analyses will be conducted using targeted sequencing panels and RNA sequencing. Integrative bioinformatic analyses will be used to identify molecular biomarkers associated with early-stage disease, advanced disease, treatment response, and therapeutic resistance. RESULTS: Ethics approval for the study has been obtained from the New Zealand Health and Disability Ethics Committee (2022 EXP 12566). This study commenced in 2022, and recruitment and biospecimen collection are ongoing. The study aims to recruit approximately 450 participants, including patients with NSCLC, individuals referred through respiratory diagnostic pathways, and nonmalignant controls. As of July 31, 2026, 205 participants have been recruited, with recruitment continuing until the target sample size is reached. Molecular and data analyses are ongoing, with additional publications expected as the cohort matures. CONCLUSIONS: This study will generate one of the first integrated genomic, epigenomic, and transcriptomic liquid biopsy datasets for NSCLC in New Zealand. The findings are expected to support the development of sensitive, accessible, and equitable blood-based biomarkers for NSCLC detection and treatment monitoring while also contributing to improved precision oncology approaches and reducing NSCLC inequities among Māori populations.

Humans

Triphenyl Phosphate Alters Methyltransferase Expression and Induces Genome-Wide Aberrant DNA Methylation in Zebrafish Larvae.

Emerging environmental contaminants, organophosphate flame retardants (OPFRs), pose significant threats to ecosystems and human health. Despite numerous studies reporting the toxic effects of OPFRs, research on their epigenetic alterations remains limited. In this study, we investigated the effects of exposure to 2-ethylhexyl diphenyl phosphate (EHDPP), tricresyl phosphate (TMPP), and triphenyl phosphate (TPHP) on DNA methylation patterns during zebrafish embryonic development. We assessed general toxicity and morphological changes, measured global DNA methylation and hydroxymethylation levels, and evaluated DNA methyltransferase (DNMT) enzyme activity, as well as mRNA expression of DNMTs and ten-eleven translocation (TET) methylcytosine dioxygenase genes. Additionally, we analyzed genome-wide methylation patterns in zebrafish larvae using reduced-representation bisulfite sequencing. Our morphological assessment revealed no general toxicity, but a statistically significant yet subtle decrease in body length following exposure to TMPP and EHDPP, along with a reduction in head height after TPHP exposure, was observed. Eye diameter and head width were unaffected by any of the OPFRs. There were no significant changes in global DNA methylation levels in any exposure group, and TMPP showed no clear effect on DNMT expression. However, EHDPP significantly decreased only DNMT1 expression, while TPHP exposure reduced the expression of several DNMT orthologues and TETs in zebrafish larvae, leading to genome-wide aberrant DNA methylation. Differential methylation occurred primarily in introns (43%) and intergenic regions (37%), with 9% and 10% occurring in exons and promoter regions, respectively. Pathway enrichment analysis of differentially methylated region-associated genes indicated that TPHP exposure enhanced several biological and molecular functions corresponding to metabolism and neurological development. KEGG enrichment analysis further revealed TPHP-mediated potential effects on several signaling pathways including TGFβ, cytokine, and insulin signaling. This study identifies specific changes in DNA methylation in zebrafish larvae after TPHP exposure and brings novel insights into the epigenetic mode of action of TPHP.

Animals

ADCY2 promoter hypomethylation in adolescents with borderline personality disorder: An RRBS study with targeted BSP replication.

BACKGROUND: Borderline Personality Disorder (BPD) is characterized by emotional dysregulation, impulsivity, and interpersonal instability. Although genetic and environmental contributions to BPD have been investigated, epigenetic correlates in adolescents remain poorly understood. The cyclic adenosine monophosphate (cAMP) signaling pathway is implicated in stress responsivity and emotion regulation, but its epigenetic variation in adolescent BPD remains understudied. METHODS: DNA methylation profiling was performed using Reduced Representation Bisulfite Sequencing (RRBS) in buccal epithelial DNA from adolescents with BPD (n = 15) and healthy controls (HC; n = 15). Differentially methylated regions (DMRs) were identified using metilene, a computational tool for detecting DMRs from bisulfite sequencing data, and subjected to Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses. Associations between methylation signals and borderline personality features were examined. Adenylate cyclase 2 (ADCY2) promoter methylation was examined using targeted bisulfite sequencing PCR (BSP) in an independent cohort (BPD n = 5; HC n = 5). RESULTS: RRBS identified 7641 DMRs between BPD and HC, with enrichment in pathways related to neuronal signaling and synaptic organization. A hypomethylated DMR was detected in the ADCY2 promoter, a gene implicated in cAMP signaling. Lower ADCY2 promoter methylation was associated with greater clinical severity, including emotional dysregulation, borderline traits, anxiety symptoms, and self-injury behaviors (r=-0.71 to -0.83; all p < 0.0001). Hypomethylation of the ADCY2 promoter was replicated in the independent cohort. CONCLUSIONS: This exploratory two-stage epigenetic study suggests that ADCY2 promoter hypomethylation in peripheral buccal epithelial DNA is associated with core symptom dimensions of adolescent BPD. These findings support a role for cAMP-related epigenetic variation in BPD and warrant replication in longitudinal studies.

ADCY2

Hybridization as driving force for cryptic species diversity in the Caribbean coral genus Madracis.

Species boundaries in scleractinian corals remain highly elusive due to conflicting patterns between morphological and molecular phylogenies, often caused by morphological plasticity, occurrence of cryptic species, incomplete lineage sorting or introgressive hybridization. Here, we use an integrated systematics approach, which combines reduced representation genome sequencing (nextRAD), micro-morphometric characterization, SEM analyses and compilation of life history traits, to infer phylogenetic relationships among closely related species in the Caribbean coral genus Madracis. In total, we analyzed 235 Madracis specimens from Cura&#xe7;ao and Bermuda collected from 10-90 m depth. Sequence- and SNP-based analyses for 115 samples generated unprecedented species resolution in Madracis, greatly supporting the morphology-based taxonomy of the current, accepted Caribbean species M. senaria, M. decactis, M. formosa, M. carmabi and M. mirabilis (M. auretenra). The exception was M. pharensis, in which we found evidence for three separate lineages, and for which we found signatures of admixture and introgression. These three M. pharensis lineages showed distinct depth distributions (thus classified as shallow, deep and very deep) and were partially distinguishable on the basis of fine microstructural elements of the collumella, septa and coenosteum. Further taxonomic comparisons are needed to formalize these putative cryptic species. Overall, our integrated systematics approach further resolves species relationships in the Caribbean genus Madracis, supports the morphological descriptions for most of the recognized species, but also reveals the existence of cryptic diversity in groups marked by high admixture, thus suggesting hybridization as a driving force in coral species diversity.

Animals

Immune Cell Type-Specific DNA Methylation Regions Associate With 24-Hour Blood Pressure Regulation in Black People.

BACKGROUND: DNA methylation and immune cells have been linked to blood pressure (BP) regulation and the development of hypertension. However, the immune cell profiles and the cell type-specific DNA methylation associated with BPs remain unclear. METHODS: This study evaluates the 19 cell type deconvolution algorithms using reduced representation bisulfite sequencing data, comparing them to in silico mixtures derived from whole-genome bisulfite sequencing. The top-performing algorithm, Epigenetic Dissection of Intra-Sample Heterogeneity (EpiDISH)-Robust Partial Correlations, was applied to 281 Black inpatients with 24-hour BP monitoring. The immune cell profiles and cell type-specific DNA methylation regions associated with these BP phenotypes were further investigated using regression analysis. RESULTS: In patients with hypertension, B-cell and CD4 effector memory T-cell abundances were significantly elevated. Monocyte and CD8 effector memory T-cell fractions positively correlated with nighttime BP, and CD3 T cells were inversely associated with office BP. These associations remained robust after covariate adjustments and were partially validated in the Medical Information Mart for Intensive Care-IV cohort. For the first time, we identified several cell type-specific DNA methylation regions as being associated with BP phenotypes and patterns across 13 immune cells, with approximately one third predominantly found in effector CD8 T cells. CONCLUSIONS: These findings provide novel insights into the epigenetically regulated immune mechanisms underlying BP regulation and identify potential targets for hypertension management.

Humans

Epigenetic safety of in vitro maturation in PCOS: genome-wide DNA methylation profiling of cord blood from a randomized controlled trial.

BACKGROUND: In vitro maturation (IVM) provides a safer alternative to conventional in vitro fertilization (IVF) for women with polycystic ovary syndrome (PCOS) by mitigating the risk of ovarian hyperstimulation. However, concerns persist regarding whether IVM perturbs epigenetic reprogramming in the offspring. Current evidence is constrained by candidate-gene approaches or a lack of parental controls. This study aimed to evaluate the genome-wide DNA methylation safety of IVM compared with conventional IVF using a rigorous trio-based design. METHODS: This secondary epigenetic analysis was nested within a randomized controlled trial (RCT) (ClinicalTrials.gov: NCT03463772). We included 10 nuclear families (trios), comprising five IVM-conceived and five IVF-conceived singleton offspring alongside their biological parents. Both groups utilized a uniform freeze-only single-blastocyst transfer strategy to minimize hormonal confounding. Genomic DNA from umbilical cord blood (UCB) and parental peripheral blood was analyzed using reduced representation bisulfite sequencing (RRBS). Genome-wide methylation patterns and differentially methylated regions (DMRs) were subsequently compared between the groups. RESULTS: Clinical characteristics were comparable between the IVM and IVF groups. Genome-wide analyses demonstrated high concordance in UCB methylation patterns, revealing no significant differences in global CpG methylation levels or distributions across key genomic features (promoters, CpG islands, and gene bodies). Only three rare DMRs were identified in UCB (representing&#x2009;~&#x2009;0.0001% of the genome), none of which mapped to imprinted or developmentally critical loci. Furthermore, methylation variability remained consistent between the groups. CONCLUSIONS: Our findings provide robust mechanistic evidence supporting the epigenetic safety of IVM. The remarkable stability of the neonatal methylome confirms that specific IVM conditions do not compromise early developmental programming, thereby endorsing IVM as a safe and viable alternative for women with PCOS. TRIAL REGISTRATION: ClinicalTrials.gov registry, NCT03463772. Registered on March 13, 2018.

Humans

A unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification.

The rapid growth of genomic sequencing demands fast, accurate, and scalable analysis methods. In viral genomic classification, expanding labeled reference collections can make supervised models costly to update and dependent on fixed label sets, motivating retrieval-based genomic classification as a simpler, more flexible alternative. We present a unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification across three viral classification tasks: hepatitis C virus (HCV) genotyping, COVID-19 discrimination, and human papillomavirus (HPV) genotyping. We compare standard sequence encodings (one-hot, k-mers, FCGR) with dense embeddings (dna2vec, DNABERT). For each representation, we evaluate supervised classifiers (Random Forest, Decision Tree, XGBoost) and retrieval-based classification, where sequence vectors are indexed with FAISS and labels are assigned via similarity-weighted k-NN. Furthermore, we benchmark multiple FAISS index types (Flat, IVF, HNSW, IVFPQ, OPQ) to characterize accuracy-speed-memory trade-offs at scale. The results show that XGBoost and retrieval using Flat or IVF indexes achieve strong classification performance under different computational profiles. Compressed indexes such as IVFPQ and OPQ substantially reduce memory usage, although their accuracy loss depends on the dataset and representation. Overall, supervised XGBoost provides a favorable accuracy-size trade-off, while retrieval-based classification remains competitive and allows labeled reference sequences to be incorporated without retraining a global classifier. This benchmark provides practical guidance for selecting sequence representations, classifiers, and vector-search indexes under different accuracy, memory, and update requirements.

Genome, Viral

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals

Genome- and peak-informed two-stage framework for scATAC-seq cell type identification.

MOTIVATION: Accurate cell type annotation is essential in scATAC-seq analysis, as it underpins the characterization of cellular heterogeneity, the identification of regulatory elements, and downstream biological discovery. However, current annotation methods still face major challenges. First, although some approaches attempt to integrate genomic sequence information, they typically rely on shallow sequence representations and thus fail to capture the long-range dependencies and regulatory signals encoded in DNA. Second, substantial batch effects introduced by different platforms, sequencing batches, or tissue sources remain insufficiently addressed. Existing models often lack robust distribution alignment and domain generalization capabilities, leading to confounding non-biological variation and reduced annotation accuracy across datasets. RESULTS: To overcome these limitations, we propose seqAlignATAC, a two-stage intra-modality annotation framework that integrates sequence-derived embeddings with domain adaptation. In the first stage, we employ a large-scale pretrained nucleotide language model to extract low-dimensional, biologically informative representations from the genomic sequences of chromatin-accessible peaks. In the second stage, these embeddings are fed into a supervised neural network equipped with an adaptive alignment module to mitigate batch effects and harmonize feature distributions between labeled reference and unlabeled target datasets. Extensive experiments across multiple settings demonstrate that seqAlignATAC achieves competitive accuracy and robustness, effectively leveraging genome-level information while alleviating batch-induced distributional discrepancies. AVAILABILITY AND IMPLEMENTATION: The source code of seqAlignATAC is available at: https://github.com/BioCS-Lab/seqAlignATAC.

Humans

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95&#xa0;% or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics

Efficient storage and regression computation for population-scale genome sequencing studies.

MOTIVATION: The growing availability of large-scale population biobanks has the potential to significantly advance our understanding of human health and disease. However, the massive computational and storage demands of whole genome sequencing (WGS) data pose serious challenges, particularly for underfunded institutions or researchers in developing countries. This disparity in resources can limit equitable access to cutting-edge genetic research. RESULTS: We present novel algorithms and regression methods that dramatically reduce both computation time and storage requirements for WGS studies, with particular attention to rare variant representation. By integrating these approaches into PLINK 2.0, we demonstrate substantial gains in efficiency without compromising analytical accuracy. In an exome-wide association analysis of 19.4 million variants for the body mass index phenotype in 125&#xa0;077 individuals (AllofUs project data), we reduced runtime from 695.35&#x2009;min (11.5&#x2009;h) on a single machine to 1.57&#x2009;min with 30 GB of memory and 50 threads (or 8.67&#x2009;min with 4 threads). Additionally, the framework supports multi-phenotype analyses, further enhancing its flexibility. AVAILABILITY AND IMPLEMENTATION: Our optimized methods are fully integrated into PLINK 2.0 and can be accessed at: https://www.cog-genomics.org/plink/2.0/.

Humans

Highly Contiguous Is Not Chromosomally Accurate: Integrated Cytogenetic and Genomic Mapping in Two Turtle Genome.

High-quality genome assemblies are essential for robust research across biological and medical fields. Assembly errors can have far-reaching consequences for downstream analyses, including gene annotation and the inference of synteny. In contrast to the rapid growth of genomic data volume, there is a notable lag in the integration of chromosome-level assemblies with cytogenetic data. We conducted the first direct genome-to-genome comparison, integrating comparative chromosome painting, the alignment of chromosome-specific probes to available genome assemblies, and synteny-based comparison of independent chromosome-level assemblies of the loggerhead sea turtle (Caretta caretta, 2n = 56) and the red-eared slider (Trachemys scripta elegans, 2n = 50). Using two independent sets of flow-sorted chromosome-specific probes in cross-species hybridizations, together with the sequencing and mapping of chromosome-derived DNA libraries, we assigned assembled scaffolds to all physical chromosomes of both species. In C. caretta, chromosomal assignments and genome-wide synteny were fully consistent with the published assembly, except for the reduced sizes of two microchromosome scaffolds, which we attribute to under-representation of repetitive DNA. In contrast, in T. s. elegans, cytogenetic validation of the assemblies revealed a false rearrangement compared to a missed one. Our results show that even highly contiguous vertebrate genome assemblies can misrepresent chromosome structure. When cytogenetic analyses reveal such inaccuracies, updated reference genomes should be generated for widely studied species to enable accurate inference of karyotype evolution and downstream comparative genomic analyses.

FISH

Equity in genome sequencing for rare disease diagnosis: a cross-sectional analysis of data from the UK 100,000 Genomes Project.

BACKGROUND: Genome sequencing has improved rare disease diagnosis and is now part of routine clinical care in the National Health Service in England. Automated prioritisation pipelines narrow millions of variants per patient to a small subset for clinical review, a process that relies on allele frequency resources that do not fully represent human genetic diversity. We assessed ancestry-related differences in variant prioritisation and diagnostic outcomes in patients from the UK 100,000 Genomes Project. METHODS: We analysed 29,405 rare disease probands with genome sequencing and linked clinical outcomes data. We used multivariable regression to assess ancestry-related differences in the number of variants prioritised for clinical review, the proportion of prioritised variants that were recorded as diagnostic, and diagnostic yield. We also evaluated the use of ancestry-stratified allele frequency filters derived from an independent, diverse UK cohort (n = 33,724). FINDINGS: Compared with the European ancestry group, the East African group had nearly three times more variants prioritised for clinical review (IRR 2.77, 95% CI 2.33-3.29). Other non-European groups also had significantly higher counts. Diagnostic yield was similar across ancestry groups after adjustment (LRT p = 0.1650). Prioritised variants were less likely to be recorded as diagnostic in East African (OR 0.32, 95% CI 0.22-0.46), West African (0.47, 0.39-0.57), South Asian (0.65, 0.58-0.73), and Middle Eastern (0.68, 0.54-0.86) groups. Applying ancestry-stratified allele-frequency filters removed 3.1% of prioritised variants overall-24.3% in the East African group-without loss of diagnostic sensitivity, including 29.5% of recorded VUS in this group. INTERPRETATION: Differences in the likelihood of prioritised variants being recorded as diagnostic partly reflect limitations of current allele frequency resources, which use broad population groupings that mask within-group diversity. Increased representation of diverse ancestries in reference databases and better estimation of ancestry-appropriate allele frequencies will help reduce inefficiencies and improve equity in variant prioritisation for rare disease diagnosis. FUNDING: The UK Department of Health and Social Care and the EU's Horizon 2020 Research and Innovation Programme.

Humans

A token-pruning framework enables efficient representation of the human genome for RNA modification analysis.

MOTIVATION: Modelling long genomic sequences remains challenging due to extreme sequence length, high redundancy, and the need for biological interpretability. Although Transformer-based architectures have achieved strong performance across genomic tasks, their high computational cost and reliance on fixed tokenization strategies limit their scalability and ability to focus on biologically informative regions. RESULTS: We propose ATSFormer, a token-pruning Transformer framework for efficient and biologically informed genomic sequence modelling. ATSFormer incorporates an attention-guided and parameter-free Adaptive Token Sampling (ATS) module into Transformer layers. Guided by attention-derived importance scores, ATS dynamically retains informative tokens while probabilistically discarding redundant ones, thereby reducing sequence length, FLOPs, and memory usage without introducing additional learnable parameters or extra training procedures. Importantly, the retained tokens correspond to key contributors to model predictions, enabling ATSFormer to highlight biologically meaningful sites and sequence motifs. We evaluated ATSFormer on four benchmark RNA modification datasets derived from RMVar 2.0, covering A-to-I, m1A, m5C, and m7G. Experimental results show that ATSFormer consistently outperforms existing state-of-the-art methods while achieving substantial computational savings. Furthermore, structural analysis using AlphaFold3 supports the biological relevance of the motifs identified by ATSFormer. AVAILABILITY AND IMPLEMENTATION: The source data and code are freely available at GitHub (https://github.com/1gao2/ATSFormer) and Zenodo (https://doi.org/10.5281/zenodo.21813541).

Humans