PubMed HealthSearch

SEARCH · PubMed Health

Results for “Overlap graph”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

7 recordsLinked to original sources

Haplotype-aware long-read error correction.

Error correction of long reads is an important initial step in genome assembly workflows. For organisms with ploidy greater than one, it is important to preserve haplotype-specific variation during read correction. This challenge has driven the development of several haplotype-aware correction methods. However, existing methods are based on either ad-hoc heuristics or deep learning approaches. In this paper, we introduce a rigorous formulation for this problem. Our approach builds on the minimum error correction framework used in reference-based haplotype phasing. We prove that the proposed formulation for error correction of reads in de novo context, i.e., without using a reference genome, is NP-hard. To make our exact algorithm scale to large datasets, we introduce practical heuristics. Experiments using PacBio HiFi sequencing datasets from human and plant genomes show that our approach achieves accuracy comparable to state-of-the-art methods. Implementation: https://github.com/at-cg/HALE .

Clustering

Rawsamble: overlapping raw nanopore signals using a hash-based seeding mechanism.

MOTIVATION: Raw nanopore signal analysis is a common approach in genomics to provide fast and resource-efficient analysis without translating the signals to bases (i.e. without basecalling). However, existing solutions cannot interpret raw signals directly if a reference genome is unknown due to a lack of accurate mechanisms to handle increased noise in pairwise raw signal comparison. Our goal is to enable the direct analysis of raw signals without a reference genome. To this end, we propose Rawsamble, the first mechanism that can identify regions of similarity between all raw signal pairs, known as all-vs-all overlapping, using a hash-based search mechanism. RESULTS: We use these overlaps to construct de novo assembly graphs with an existing assembler, miniasm, off-the-shelf. To our knowledge, these are the first de novo assemblies ever constructed directly from raw signals without basecalling. Our extensive evaluations across multiple genomes of varying sizes show that Rawsamble provides a significant speedup (on average by 5.01× and up to 23.10×) and reduces peak memory usage (on average by 5.74× and up to by 22.00×) compared to a conventional genome assembly pipeline using the state-of-the-art tools for basecalling (Dorado's fastest mode) and overlapping (minimap2) on a CPU. We find that around one-third of Rawsamble's overlapping pairs are also found by minimap2. We find that when we use overlapping reads from Rawsamble, we can construct unitigs that are (i) as accurate as those built from minimap2's overlaps and (ii) up to half a chromosome in length (e.g. 2.3 million bases for E. coli). AVAILABILITY AND IMPLEMENTATION: Rawsamble is available at https://github.com/CMU-SAFARI/RawHash. We also provide the scripts to fully reproduce our results on our GitHub page.

Nanopores

Autocycler: long-read consensus assembly for bacterial genomes.

MOTIVATION: Long-read sequencing enables complete bacterial genome assemblies, but individual assemblers are imperfect and often produce sequence-level and structural errors. Consensus assembly using Trycycler can improve accuracy, but its lack of automation limits scalability. There is a need for an automated method to generate high-quality consensus bacterial genome assemblies from long-read data. RESULTS: We present Autocycler, a command-line tool for generating accurate bacterial genome assemblies by combining multiple alternative long-read assemblies of the same genome. Without requiring user input, Autocycler builds a compacted De Bruijn graph from the input assemblies, clusters and filters contigs, trims overlaps, and resolves consensus sequences by selecting the most common variant at each locus. It also supports manual curation when desired, allowing users to refine assemblies in challenging or important cases. In our evaluation using Oxford Nanopore Technologies reads from five bacterial isolates, Autocycler outperformed individual assemblers, automated pipelines, and other consensus tools, producing assemblies with lower error rates and improved structural accuracy. AVAILABILITY AND IMPLEMENTATION: Autocycler is implemented in Rust, open-source, and freely available at github.com/rrwick/Autocycler. It runs on Linux and macOS and is extensively documented.

Genome, Bacterial

plinkQC: an integrated tool for ancestry inference, sample selection, and quality control in population genetics.

MOTIVATION: Population genetic analyses rely on high quality datasets that pass rigorous controls for sample and marker quality. Many analyses also require additional processing including identification of ancestry and sample relatedness. A software package that addresses all these common, yet crucial tasks is missing. RESULTS: We have developed plinkQC, an R/CRAN package that combines these functionalities into a single software package with detailed vignettes for example applications. plinkQC determines the ancestry of study samples via a pre-trained random forest classifier that reaches 98% performance accuracy with just 5% of marker overlap between reference and user data. To obtain the maximal set of unrelated study samples, we developed a graph-based pruning method, taking both relationship estimates and sample quality into account. We demonstrate optimal sample selection on the 1000 Genomes project, where we retain an additional 71 samples compared to publicly available exclusion lists. Finally, plinkQC bundles these results together with per-individual and per-marker quality control checks into three simple functions and returns both the quality controlled dataset and quality control report about each step of the analysis. AVAILABILITY AND IMPLEMENTATION: plinkQC is available as an R/CRAN package. The documentation and code are available on github: https://meyer-lab-cshl.github.io/plinkQC/ and https://github.com/meyer-lab-cshl/plinkQC_manuscript.

Software

Artificial Intelligence-Driven Multi-Omics Analysis Reveals Hydroxytyrosol Targeting of the TXNIP-NLRP3 Inflammasome Axis in Traumatic Brain Injury.

Traumatic brain injury (TBI) induces secondary neuroinflammation driven by oxidative stress, inflammasome activation, and immune remodeling, yet specific mechanism-guided pharmacological interventions remain limited. This study established an artificial intelligence (AI)-integrated network pharmacology and multi-omics framework to evaluate whether hydroxytyrosol (HT), an olive-derived natural polyphenol, may regulate TBI-related neuroinflammatory targets centered on the TXNIP/NLRP3 inflammasome axis. Starting from the SMILES structure of HT, potential targets were predicted using PharmMapper, SwissTargetPrediction, and the Similarity Ensemble Approach and were standardized to UniProt identifiers. TBI-associated genes were integrated from GeneCards, DisGeNET, OMIM, and the Therapeutic Target Database. The overlapping target set was analyzed using STRING-based protein-protein interaction (PPI) networks, MCODE, CytoHubba, Gene Ontology (GO), and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment. Public GEO transcriptomic datasets (GSE123831 and GSE104687) were used for cross-platform expression validation, differential expression analysis, and exploratory CIBERSORT-based immune infiltration estimation. Random forest (RF), multilayer perceptron (MLP), graph convolutional network (GCN), graph attention network (GAT), SHAP/LIME explainability analysis, LASSO inflammatory-risk scoring, and two-sample Mendelian randomization (MR) were further applied for target prioritization, immune phenotype mapping, and genetic association analysis. Seventy-three overlapping HT-TBI targets were identified. PPI and topology analyses prioritized TXNIP, NLRP3, CASP1, MAPK1, and TP53 as key hubs enriched in inflammasome activation, oxidative stress, apoptosis, and NOD-like receptor signaling. TXNIP, NLRP3, and CASP1 were consistently upregulated in both TBI transcriptomic datasets. LM22-based immune deconvolution suggested increased pro-inflammatory immune signatures and a positive TXNIP-M1 macrophage association (r&#x202f;=&#x202f;0.63, p < 0.001), which should be interpreted as a transcriptome-derived hypothesis rather than validated murine immune-cell proportions. AI-based models consistently ranked TXNIP/NLRP3 as high-contribution features under internal validation, and removal of these targets reduced model performance. A five-gene inflammatory score achieved an internally evaluated AUC of 0.87, while two-sample MR supported positive genetic associations involving TXNIP expression, TBI risk, NLRP3 and IL-1&#x3b2; expression. Collectively, these findings prioritize the TXNIP/NLRP3/CASP1 module as a computationally supported candidate mechanism through which HT may influence oxidative stress-inflammasome-immune coupling in TBI. This study provides an interpretable drug-target-pathway-phenotype framework and identifies TXNIP, NLRP3, and CASP1 as priority nodes for future experimental validation.

Artificial Intelligence

Optimizing sparse and skew hashing: faster k-mer dictionaries.

MOTIVATION: Representing a set of k-mers-strings of length k-in small space under fast lookup queries is a fundamental requirement for several applications in Bioinformatics. A data structure based on sparse and skew hashing (SSHash) was recently proposed for this purpose (Pibiri 2022): it combines good space effectiveness with fast lookup and streaming queries. It is also order-preserving, i.e. consecutive k-mers (sharing a prefix-suffix overlap of length k-1) are assigned consecutive hash codes which helps compressing satellite data typically associated with k-mers, like abundances and color sets in colored De Bruijn graphs. RESULTS: We study the problem of accelerating queries under the sparse and skew hashing indexing paradigm, without compromising its space effectiveness. We propose a refined data structure with less complex lookups and fewer cache misses. We give a simpler and faster algorithm for streaming lookup queries. The refined architecture translates to substantial performance gains, outperforming the original version of SSHash in both index construction speed and query efficiency. Compared to indexes with similar capabilities and based on the Burrows-Wheeler transform, like SBWT and FMSI, SSHash is significantly faster to build and query. SSHash is competitive in space with the fast (and default) modality of SBWT when both k-mer strands are indexed. While larger than FMSI, it is also more than one order of magnitude faster to query. AVAILABILITY AND IMPLEMENTATION: The SSHash software is available at https://github.com/jermp/sshash, and also distributed via Bioconda. A benchmark of data structures for k-mer sets is available at https://github.com/jermp/kmer_sets_benchmark. The datasets used in this article are described and available at https://zenodo.org/records/17582116.

Algorithms

Robust error-minimization in the genetic code across physicochemical metrics and variant codes: A graph-theoretic analysis in GF(2)6.

The standard genetic code reduces the impact of point mutations, but the robustness of this property across physicochemical metrics, naturally occurring variant codes, and codon-reassignment mechanisms remains incompletely quantified. Embedding the 64 codons in GF(2)6 represents the hypercube Q6 as a coordinate-dependent subgraph of the encoding-independent single-nucleotide mutation graph H(3,4), and enables continuous &#x3c1;-interpolation between the two. Under a quartet-pattern shuffle null (n=10,000), the standard code is significantly low-cost across four established, code-independent physicochemical distance metrics with partially overlapping content (Grant ham p=0.0062; Miyata p<0.001; Woese polar requirement p=0.003; Kyte-Doolittle hydropathy p=0.001), and the signal strengthens monotonically as &#x3c1; moves Q6&#x2192;H(3,4). A structure-aware sensitivity analysis under the alignment-derived ProtSub matrix (Jia & Jernigan 2021) yields the most extreme percentile of any measure tested (p=0.0004; all five p-values pass Bonferroni at &#x3b1;=0.05). Across the 27 NCBI translation tables, near-optimality is preserved: 11 of 12 informative-distance variants retain top-5% placement after BH-FDR correction. Natural codon reassignments avoid disrupting codon-family connectivity: under the encoding-independent H(3,4) adjacency, observed events are topology-breaking at relative risk 0.32 versus the candidate landscape (permutation p&#x2264;10-4). The H(3,4) result is stable by construction; the Q6 decomposition is representation-specific and fails to show depletion under 8 of 24 base-to-bit encodings, so we report H(3,4) as the primary test and Q6 as a sensitivity. Event-level conditional-logit modelling shows that topology avoidance and local physicochemical cost provide complementary, only weakly correlated signal (rs=0.15), and that topology adds explanatory value beyond physicochemistry under both Q6 and encoding-independent H(3,4) adjacency. Retrospective reanalysis of nine genome-recoding datasets is consistent with codon-family topology operating as an evolutionary-trajectory constraint distinct from acute engineering fitness. The contribution is the second axis: code evolution is jointly constrained by physicochemical smoothness and codon-family topological integrity, and these two constraints are partly independent.

Codon reassignment