PubMed HealthSearch

SEARCH · PubMed Health

Results for “reference-free”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

15 recordsLinked to original sources

Reference-Free Variant Calling with Local Graph Construction with ska lo (SKA).

The study of genomic variants is increasingly important for public health surveillance of pathogens. Traditional variant-calling methods from whole-genome sequencing data rely on reference-based alignment, which can introduce biases and require significant computational resources. Alignment- and reference-free approaches offer an alternative by leveraging k-mer-based methods, but existing implementations often suffer from sensitivity limitations, particularly in high mutation density genomic regions. Here, we present ska lo, a graph-based algorithm that aims to identify within-strain variants in pathogen whole-genome sequencing data by traversing a colored De Bruijn graph and building variant groups (i.e. sets of variant combinations). Through in silico benchmarking and real-world dataset analyses, we demonstrate that ska lo achieves high sensitivity in single-nucleotide polymorphism (SNP) calls while also enabling the detection of insertions and deletions, as well as SNP positioning on a reference genome for recombination analyses. These findings highlight ska lo as a simple, fast, and effective tool for pathogen genomic epidemiology, extending the range of reference-free variant-calling approaches. ska lo is freely available as part of the SKA program (https://github.com/bacpop/ska.rust).

Polymorphism, Single Nucleotide

CamK-DB: A k-mer MinHash fingerprint database for reference-free genotyping of Camellia accessions.

Tea (Camellia sinensis L.), a major global economic crop in Asia, poses challenges for genetic identification because its highly heterozygous, repetitive genome reduces the efficacy of conventional single-nucleotide polymorphism (SNP) and microsatellite markers, and interspecific hybridization further complicates the situation. To address these issues, CamK-DB was developed as a reference-free Camellia fingerprinting database built on MIKE MinHash sketches. We curated 418 candidate resequencing datasets, and built a database using standardized 5× genome-coverage fingerprints. Each accession is stored as a MIKE. jac fingerprint generated with k = 21 and recommended sketch/pre_cnt = 2000. CamK-DB provides a command-line interface for data management and a custom C++ query engine that computes top-10 matches using Jaccard similarity, complemented by a QT-based graphical interface for interactive analysis. This resource offers a robust and scalable framework for precise and routine germplasm identification, genomic phylogenetic inference, and strategic breeding program design. CamK-DB (database and code) is publicly available at https://github.com/sc-zhang/CamK-DB. CamK-DB binaries are provided for Windows 10/11 and Linux (x86_64, glibc ≥ 2.27).

Databases, Genetic

Reference-Free Microsatellite Instability Detection from Tumor Sequencing Using Intrasample Variability Modeling.

Microsatellite instability (MSI) is a predictive biomarker in several tumor types. However, many next-generation sequencing-based callers require matched normal samples, reference panels, or pretrained models, limiting their portability across assays and sequencing centers. We developed PROMIS (PROfiling of Microsatellite InStability), a tumor-only, reference-free pipeline that uses a discrete mixture model to characterize intrasample repeat-length distributions at predefined microsatellite loci. Locus-level classifications are then aggregated into a continuous MSI score. We benchmarked PROMIS in colorectal (CRC), endometrial (UCEC), and gastric (STAD) cancers from The Cancer Genome Atlas. PROMIS achieved an overall area under the receiver operating characteristic curve (AUC) of 0.995 and cohort-specific AUCs of 1.00 in CRC and stomach adenocarcinoma and 0.999 in uterine corpus endometrial carcinoma, comparable to established tools despite not using matched normals or pretrained models. Subsampling demonstrated robust performance with substantially fewer loci. In silico dilution showed progressively reduced MSI-microsatellite-stable discrimination, with the pooled AUC declining from 0.83 at 10% tumor fraction to 0.53 at 1%. At low tumor fractions, tumor-type-specific baseline microsatellite variability increasingly influenced PROMIS scores. Finally, in prostate and CRC cell-free DNA cohorts, including Illumina TSO500 data and an 18-gene panel, PROMIS yielded MSI scores concordant with orthogonal tissue- and panel-based classifications across the evaluated Illumina-based sequencing contexts. Accordingly, the present validation should be considered limited to Illumina-based sequencing platforms. PROMIS is intended to complement existing genomic profiling workflows by enabling MSI assessment from sequencing data already generated for broader molecular analyses. Prospective clinical validation remains necessary before clinical implementation.

Journal Article

MKMC enables reference-free transcriptomic analysis using k-mer representations.

Traditional RNA-seq analysis depends heavily on genome alignment and gene annotation, limiting its utility in non-model organisms and introducing biases that can obscure regulatory complexity. We present MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer-based statistics to detect biological variation without requiring alignment. MKMC integrates fast k-mer counting, abundance matrix generation, normalization, dimensionality reduction, and differential analysis into a unified workflow. Across diverse datasets, MKMC recapitulates key biological signals-including sex differences in killifish liver-and matches alignment-based pipelines in differential expression analysis and transcriptomic age prediction. Notably, MKMC detects isoform-specific events missed by traditional methods, one of which we validated using in situ hybridization. These results reveal previously hidden isoform-level regulatory events that contribute to sex- and age-associated transcriptional programs. MKMC offers a robust, extensible alternative to alignment-based approaches, enabling transcriptomic discovery across both model and non-model systems. While we focus here on RNA-seq as a primary application, MKMC is broadly applicable to any k-mer-based analysis of next-generation sequencing data.

MKMC

UMI-nea: a fast, robust tool for reference-free UMI deduplication and accurate quantification.

MOTIVATION: One of the key applications of Unique Molecular Identifiers (UMIs) in high-throughput sequencing is to correct for PCR amplification bias and removal of PCR duplicates, thereby improving quantification in DNA-seq and RNA-seq applications. Accurately grouping error-bearing UMIs that originate from the same input molecule through a UMI deduplication method is a critical step in this process. However, many existing UMI deduplication tools rely on simple Hamming distance comparisons or suboptimal clustering algorithms, often resulting in erroneous UMI groupings, particularly in error-prone long-read sequencing or ultra-high-depth short-read sequencing. RESULTS: We introduce UMI-nea, a tool that utilizes Levenshtein distance comparisons and a novel clustering approach to optimize multithreading workflows. Compared against three other indel-aware UMI deduplication tools, UMI-nea achieves more accurate UMI groupings with efficient run time. It demonstrates robust performance across diverse sequencing platforms, depths, and UMI lengths. Additionally, UMI-nea incorporates a data-guided adaptive UMI filter, further enhancing quantification accuracy. AVAILABILITY AND IMPLEMENTATION: UMI-nea is available on github https://github.com/Qiaseq-research/UMI-nea.git or Zenodo https://doi.org/10.5281/zenodo.16745758. Sequencing data are stored at https://qiagenpublic.blob.core.windows.net/umi-nea-datasets/.

High-Throughput Nucleotide Sequencing

An archaic reference-free method to jointly infer Neanderthal and Denisovan introgressed segments in modern human genomes.

Admixture between populations is a common feature of human history. Admixture events introduce new genetic variation that can fuel evolution. Characterizing the significance of admixture events on the evolution of populations across various species is of great interest to evolutionary geneticists. Local Ancestry Inference (LAI) methods infer genetic ancestry of an individual at a particular chromosomal location. Certain methods specialize in detecting archaic introgression, which consists of interbreeding between modern and archaic humans like Neanderthals and Denisovans. Most current LAI methods allow the detection of a single archaic ancestry, and post-processing may distinguish between multiple waves of introgression. These methods vary in how they choose archaic or modern reference genomes for the inference. Here, we present a new HMM-based method (DAIseg), which has the advantage of simultaneously distinguishing between multiple waves of ancient and recent admixture, using only modern human reference genomes. Simulations demonstrate that DAIseg achieves higher overall performance than state-of-the-art methods. We also apply DAIseg to Papuan populations to jointly detect Denisovan and Neanderthal introgressed segments, and identify a higher number of archaic segments than previous methods. Analysis of inferred introgressed segments, shows that we can identify evidence for two Denisovan introgression events in Papuans. Overall, on top of being able to deal with both Archaic and recent admixture, DAIseg provides a more principled approach for detecting and classifying Denisovan and Neanderthal segments which will improve downstream analysis of introgressed segments to infer the impact of archaic introgression in humans.

Denisovan

Embed-Search-Align: DNA sequence alignment using Transformer models.

MOTIVATION: DNA sequence alignment, an important genomic task, involves assigning short DNA reads to the most probable locations on an extensive reference genome. Conventional methods tackle this challenge in two steps: genome indexing followed by efficient search to locate likely positions for given reads. Building on the success of Large Language Models in encoding text into embeddings, where the distance metric captures semantic similarity, recent efforts have encoded DNA sequences into vectors using Transformers and have shown promising results in tasks involving classification of short DNA sequences. Performance at sequence classification tasks does not, however, guarantee sequence alignment, where it is necessary to conduct a genome-wide search to align every read successfully, a significantly longer-range task by comparison. RESULTS: We bridge this gap by developing a "Embed-Search-Align" (ESA) framework, where a novel Reference-Free DNA Embedding (RDE) Transformer model generates vector embeddings of reads and fragments of the reference in a shared vector space; read-fragment distance metric is then used as a surrogate for sequence similarity. ESA introduces: (i) Contrastive loss for self-supervised training of DNA sequence representations, facilitating rich reference-free, sequence-level embeddings, and (ii) a DNA vector store to enable search across fragments on a global scale. RDE is 99% accurate when aligning 250-length reads onto a human reference genome of 3 gigabases (single-haploid), rivaling conventional algorithmic sequence alignment methods such as Bowtie and BWA-Mem. RDE far exceeds the performance of six recent DNA-Transformer model baselines such as Nucleotide Transformer, Hyena-DNA, and shows task transfer across chromosomes and species. AVAILABILITY AND IMPLEMENTATION: Please see https://anonymous.4open.science/r/dna2vec-7E4E/readme.md.

Sequence Analysis, DNA

De novo clustering of large long-read transcriptome datasets with isONclust3.

MOTIVATION: Long-read sequencing techniques can sequence transcripts from end to end, greatly improving our ability to study the transcription process. Although there are several well-established tools for long-read transcriptome analysis, most are reference-based. This limits the analysis of organisms without high-quality reference genomes and samples or genes with high variability (e.g. cancer samples or some gene families). In such settings, analysis using a reference-free method is favorable. The computational problem of clustering long reads by region of common origin is well-established for reference-free transcriptome analysis pipelines. Such clustering enables large datasets to be split roughly by gene family and, therefore, an independent analysis of each cluster. There exist tools for this. However, none of those tools can efficiently process the large amount of reads that are now generated by long-read sequencing technologies. RESULTS: We present isONclust3, an improved algorithm over isONclust and isONclust2, to cluster massive long-read transcriptome datasets into gene families. Like isONclust, isONclust3 represents each cluster with a set of minimizers. However, unlike other approaches, isONclust3 dynamically updates the cluster representation during clustering by adding high-confidence minimizers from new reads assigned to the cluster and employs an iterative cluster-merging step. We show that isONclust3 yields results with higher or comparable quality to state-of-the-art algorithms but is 10-100 times faster on large datasets. Also, using a 256 Gb computing node, isONclust3 was the only tool that could cluster 37 million PacBio reads, which is a typical throughput of the recent PacBio Revio sequencing machine. AVAILABILITY AND IMPLEMENTATION: https://github.com/aljpetri/isONclust3.

Algorithms

Artifactually high coherences result from using spherical spline computation of scalp current density.

Coherence computed from common reference montages inextricably confounds true coherence with power and phase at the recording and reference electrodes. Direct measurement of coherence requires reference-free EEG data, such as data from EEG scalp current densities (SCDs), which estimate the potential gradient perpendicular to the scalp. Perrin et al. (1989) presented a method for computing SCDs by taking the Laplacian of the scalp potential surface generated by spherical spline interpolation. When this method of computing SCDs was applied to EEG data gathered from young adults, very high values were observed for inter-electrode coherences computed from the spherical spline derived SCD data but not from coherences computed from the common reference data. These high coherences prompted further examination of the properties of the spherical spline function and of spherical spline derived SCDs. Simulated data were constructed, and coherence was computed on the simulated data and on the SCDs derived from the spherical spline procedure and from the Hjorth (1980) procedure. The results of those simulations are presented, which demonstrate that a major artifact is introduced by using the spherical spline procedure. This artifact results from the spline weighting matrix used to derive the SCDs and strongly inflates the inter-electrode coherences of the SCD transformed data.

Artifacts

Demixer: a probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample.

MOTIVATION: Multi-drug resistant or hetero-resistant tuberculosis (TB) hinders the successful treatment of TB. Hetero-resistant TB occurs when multiple strains of the TB-causing bacterium with varying degrees of drug susceptibility are present in an individual. Existing studies predicting the proportion and identity of strains in a mixed infection sample rely on a reference database of known strains. A main challenge then is to identify de novo strains not present in the reference database, while quantifying the proportion of known strains. RESULTS: We present Demixer, a probabilistic generative model that uses a combination of reference-based and reference-free techniques to delineate mixed infection strains in whole genome sequencing (WGS) data. Demixer extends a topic model widely used in text mining to represent known mutations and discover novel ones. Parallelization and other heuristics enabled Demixer to process large datasets like CRyPTIC (Comprehensive Resistance Prediction for Tuberculosis: an International Consortium). In both synthetic and experimental benchmark datasets, our proposed method precisely detected the identity (e.g. 91.67% accuracy on the experimental in vitro dataset) as well as the proportions of the mixed strains. In real-world applications, Demixer revealed novel high confidence mixed infections (101 out of 1963 Malawi samples analysed), and new insights into the global frequency of mixed infection (2% at the most stringent threshold in the CRyPTIC dataset) and its significant association to drug resistance. Our approach is generalizable and hence applicable to any bacterial and viral WGS data. AVAILABILITY AND IMPLEMENTATION: All code relevant to Demixer is available at https://github.com/BIRDSgroup/Demixer.

Mycobacterium tuberculosis

Genome size estimation from long read overlaps.

MOTIVATION: Accurate genome size estimation is an important component of genomic analyses such as assembly and coverage calculation, though existing tools are primarily optimized for short-read data. RESULTS: We present LRGE, a novel tool that uses read-to-read overlap information to estimate genome size in a reference-free manner. LRGE calculates per-read genome size estimates by analysing the expected number of overlaps for each read, considering read lengths and a minimum overlap threshold. The final size is taken as the median of these estimates, ensuring robustness to outliers such as reads with no overlaps. Additionally, LRGE provides an expected confidence range for the estimate. We validate LRGE on a large, diverse bacterial dataset and confirm it generalizes to eukaryotic datasets. On bacterial genomes, LRGE outperforms k-mer-based methods in both accuracy and computational efficiency and produces genome size estimates comparable to those from assembly-based approaches, like Raven, while using significantly less computational resources. AVAILABILITY AND IMPLEMENTATION: Our method, LRGE (Long Read-based Genome size Estimation from overlaps), is implemented in Rust and is available as a precompiled binary for most architectures, a Bioconda package, a prebuilt container image, and a crates.io package as a binary (lrge) or library (liblrge). The source code is available at https://github.com/mbhall88/lrge and an archive at https://doi.org/10.5281/zenodo.17183812 under an MIT license.

Genome Size

MetaFX: feature extraction from whole-genome metagenomic sequencing data.

MOTIVATION: Microbial communities consist of thousands of microorganisms and viruses and have a tight connection with an environment, such as gut microbiota modulation of host body metabolism. However, the direct relationship between the presence of certain microorganism and the host state often remains unknown. Toolkits using reference-based approaches are limited to microbes present in databases. Reference-free methods often require enormous resources for metagenomic assembly or results in many poorly interpretable features based on k-mers. RESULTS: Here we present MetaFX-an open-source library for feature extraction from whole-genome metagenomic sequencing data and classification of groups of samples. Using a large volume of metagenomic samples deposited in databases, MetaFX compares samples grouped by metadata criteria (e.g. disease, treatment, etc.) and constructs genomic features distinct for certain types of communities. Features constructed based on statistical k-mer analysis and de Bruijn graphs partition. Those features are used in machine learning models for classification of novel samples. Extracted features can be visualized on de Bruijn graphs and annotated for providing biological insights. We demonstrate the utility of MetaFX by building classification models for 590 human gut samples with inflammatory bowel disease. Our results outperform the previous research disease prediction accuracy up to 17%, and improves classification results compared to taxonomic analysis by 9±10% on average. AVAILABILITY AND IMPLEMENTATION: MetaFX is a feature extraction toolkit applicable for metagenomic datasets analysis and samples classification. The source code, test data, and relevant information for MetaFX are freely accessible at https://github.com/ctlab/metafx under the MIT License. Alternatively, MetaFX can be obtained via http://doi.org/10.5281/zenodo.16949369.

Metagenomics

MetaflowX: a scalable and resource-efficient workflow for multi-strategy metagenomic analysis.

Microbiomes play crucial roles in diverse ecosystems, spanning environmental, agricultural, and human health domains. However, in-depth metagenomic data analysis presents significant technical and resource challenges, particularly at scale. Existing computational pipelines are typically limited to either reference-based or reference-free approaches and exhibit inefficiencies in process large datasets. Here, we introduce MetaflowX (https://github.com/01life/MetaflowX), an open-resource workflow integrating both analytical paradigms for enhanced metagenomic investigations. This modular framework encompasses short-read quality control, rapid microbial profiling, hybrid contig assembly and binning, high-quality metagenome-assembled genome (MAG) identification, as well as bin refinement and reassembly. Benchmarking tests showed that MetaflowX completed full metagenomic analyses up to 14-fold faster and with 38% less disk usage than existing workflows. It also recovered the highest number of high-quality and taxonomically diverse MAGs. A dedicated reassembly module further improved MAG quality, increasing completeness by 5.6% and reducing contamination by 53% on average. Functional annotation modules enable detection of key features, including virulence and antibiotic resistance genes. Designed for extensibility, MetaflowX provides an efficient solution addressing current and emerging demands in large-scale metagenomic research.

Metagenomics

Diversification, loss, and virulence gains of the major effector AvrStb6 during continental spread of the wheat pathogen Zymoseptoria tritici.

Interactions between plant pathogens and their hosts are highly dynamic and mainly driven by pathogen effectors and plant receptors. Host-pathogen co-evolution can cause rapid diversification or loss of pathogen genes encoding host-exposed proteins. The molecular mechanisms that underpin such sequence dynamics remains poorly investigated at the scale of entire pathogen species. Here, we focus on AvrStb6, a major effector of the global wheat pathogen Zymoseptoria tritici, evolving in response to the cognate receptor Stb6, a resistance widely deployed in wheat. We comprehensively captured effector gene evolution by analyzing a global thousand-genome panel using reference-free sequence analyses. We found that AvrStb6 has diversified into 59 protein isoforms with a strong association to the pathogen spreading to new continents. Across Europe, we found the strongest differentiation of the effector consistent with high rates of Stb6 deployment. The AvrStb6 locus showed also a remarkable diversification in transposable element content with specific expansion patterns across the globe. We detected AvrStb6 gene losses and evidence for transposable element-mediated disruptions. We used virulence datasets of genome-wide association mapping studies to predict virulence changes across the global panel. Genomic predictions suggested marked increases in virulence on Stb6 cultivars concomitant with the spread of the pathogen to Europe and the subsequent spread to further continents. Finally, we genotyped French bread wheat cultivars for Stb6 and monitored resistant cultivar deployment concomitant with AvrStb6 evolution. Taken together, our data provides a comprehensive view of how a rapidly diversifying effector locus can undergo large-scale sequence changes concomitant with gains in virulence on resistant cultivars. The analyses highlight also the need for large-scale pathogen sequencing panels to assess the durability of resistance genes and improve the sustainability of deployment strategies.

Ascomycota

Three-dimensional reconstruction of single particles embedded in ice.

Single particles embedded in ice pose new challenges for image processing because of the intrinsically low signal-to-noise ratio of such particles in electron micrographs. We have developed new techniques that address some of these problems and have applied these techniques to electron micrographs of the Escherichia coli ribosome. Data collection and reconstruction follow the protocol of the random-conical technique of Radermacher et al. [J. Microscopy 146 (1987) 113]. A reference-free alignment algorithm has been developed to overcome the propensity of reference-based algorithms to reinforce the reference motif in very noisy situations. In addition, an iterative 3D reconstruction method based on a chi-square minimization constraint has been developed and tested. This algorithm tends to reduce the effects of the missing angular range on the reconstruction, thereby facilitating the merging of random-conical data sets obtained from differently oriented particles.

Algorithms