PubMed HealthSearch

SEARCH · PubMed Health

Results for “Performance benchmarking”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Pangenomes aid accurate detection of large insertions and deletions from targeted sequencing: the case of cardiomyopathies.

BACKGROUND: Gene panels represent a widely used strategy for genetic testing in a vast range of Mendelian disorders. While this approach aids reliable bioinformatic detection of short coding variants, it often fails to detect many larger variants. Recent studies have recommended the adoption of pangenome references (as opposed to linear reference genomes like GRCh38) to augment detection of large variants from targeted sequencing, potentially providing diagnostic laboratories with the possibility to streamline diagnostic work-ups and reduce costs. METHODS: Here, we analyze 1969 cardiomyopathy cases and 1805 controls sequenced with the Illumina Trusight Cardio panel using a pangenome-based workflow (GRAF) and five conventional orthogonal methodologies (GATK HaplotypeCaller, GATK-gCNV, ExomeDepth, Manta and Lumpy-SV) to detect variants ≥ 20 bp in size. RESULTS: Following lab-based variant validation by means of PCR and Sanger sequencing, we show that GRAF conjugates higher precision and recall (F1 score 0.86) compared with other methods (F1 0-0.57) in detecting potentially pathogenic variants ≥ 20 bp from short-read panel data. Results were complemented by a comparison of the tools' performance in detecting ground truth variants on reference sample HG002 from Genome In A Bottle, which confirmed GRAF to outperform other tools also on exome sequencing (F1 0.97 vs. 0-0.94). Notably, in the HG002 benchmark dataset, GRAF also showed slightly improved performance compared to GATK HaplotypeCaller in the identification of small variants (1-19 bp; F1 0.975 vs. 0.968). CONCLUSIONS: Our results indicate that pangenome-based workflows aid improved detection of large variants from targeted sequencing data in the clinical context and suggest that they may contribute to more unified variant detection frameworks for all-size genetic variants in the future.

Humans

Evaluation of fellowship and residency programs in a comprehensive cancer center.

A Dental Oncology residency and fellowship program was evaluated annually from 1980-1984 by each trainee. Nineteen training program variables were assessed on an Exit from Training Questionnaire using 0-100 mm linear scales. The five annual scores for each variable were subjected to trend analysis. Implication of the results to the training program are discussed. The evaluation method may be used to establish a departmental training program benchmark in relation to which future performance can be measured and as a guide to program modification. The method may be applied to all departmental training programs and, when combined, to the institution as a whole.

Cancer Care Facilities

Meta-analysis models with group structure for pleiotropy detection at gene and variant level using summary statistics from multiple datasets.

Genome-wide association studies (GWASs) have highlighted the importance of pleiotropy in human diseases, where one gene can impact 2 or more unrelated traits. Examining shared genetic risk factors across multiple diseases can enhance our understanding of these conditions by pinpointing new genes and biological pathways involved. Furthermore, with an increasing wealth of GWAS summary statistics available to the scientific community, leveraging these findings across multiple phenotypes could unveil novel pleiotropic associations. Existing selection methods examine pleiotropic associations one by one at a scale of either the genetic variant or the gene, and thus cannot consider all the genetic information at the same time. To address this limitation, we propose a new approach called MPSG (Meta-analysis model adapted for Pleiotropy Selection with Group structure). This method performs a penalized multivariate meta-analysis method adapted for pleiotropy and takes into account the group structure information nested in the data to select relevant variants and genes (or pathways) from all the genetic information. To do so, we implemented an alternating direction method of multipliers algorithm. We compared the performance of the method with other benchmark meta-analysis approaches such as GCPBayes, PLACO, and ASSET by considering as inputs different kinds of summary statistics. We provide an application of our method to the identification of potential pleiotropic genes between breast and thyroid cancers.

Humans

Identifying features, performance, and limitations of dairy ration formulation software: a comparison of three ration formulation programs.

A method for evaluating and selecting dairy ration formulation software for microcomputers is presented. Information obtained from a survey of practicing nutritionists and from interviews with computer programmers was used to identify users' needs, desired features, and patterns of use. A benchmark problem consisting of 15 activities frequently required in designing dairy rations was developed to evaluate ease of use, ease of learning, and software performance. Ease of use was measured by counting keystrokes and recording time required by a user familiar with these ration formulation programs to complete the benchmark problem. Ease of learning was measured as the amount of help and time needed by users unfamiliar with this software to complete the benchmark problem. Software performance was measured by speed of formulation, ration costs, ingredients and amounts selected, and nutrient content of rations. Evaluation of three commercially available software applications was made using the method developed.

Animal Feed

Resolution-dependent self-supervised transfer in chest radiograph classification.

BACKGROUND: Self-supervised learning (SSL) has improved visual representation learning, but its value in chest radiography remains uncertain. DINOv3 extends earlier SSL models through Gram-anchored self-distillation and explicit high-resolution adaptation. Whether these changes improve transfer learning for chest radiograph classification has not been established. METHODS: We benchmarked DINOv3 against DINOv2 and supervised ImageNet initialization across seven chest radiograph datasets comprising 816,183 radiographs from pediatric and adult cohorts. ViT-B/16 and ConvNeXt-B were evaluated under full fine-tuning at 224 × 224 and 512 × 512 pixels, with targeted 1024 × 1024 experiments on three cohorts. Additional analyses examined parameter-efficient adaptation, synthetic label corruption, external validation, frozen 7B features, and computational efficiency. The primary outcome was the mean area under the receiver operating characteristic curve across labels. RESULTS: In adult cohorts, DINOv3 did not consistently outperform DINOv2 at 224 × 224 pixels, but became the strongest initialization at 512 × 512 pixels, especially with ConvNeXt-B. Gains were greatest for small focal and boundary-dependent abnormalities, whereas large-structure findings changed little. The pediatric cohort showed no significant benefit from DINOv3, higher resolution, or backbone choice. Scaling to 1024 × 1024 rarely improved performance and markedly increased computational cost. ConvNeXt-B remained superior to ViT-B/16 under both full and parameter-efficient adaptation. External validation preserved the 512 × 512 DINOv3 advantage, whereas synthetic label corruption showed that this benefit should not be interpreted simply as superior noise robustness. Frozen DINOv3-7B features underperformed relative to fully adapted 86 to 89M-parameter backbones. CONCLUSIONS: For adult chest radiograph classification, DINOv3 provides its most reliable benefit at 512 × 512 pixels, particularly with ConvNeXt-B. Fully adapted mid-sized models at 512 × 512 pixels provided the best performance-cost trade-off in our benchmark.

Journal Article

Variant harmonization critically determines polygenic score transferability for lipid traits in Samoan populations.

Dyslipidemia is a significant risk factor for cardiovascular disease (CVD), the leading cause of death in Samoa. Polygenic scores (PGSs) for lipid traits offer promise for improved CVD risk prediction; however, their performance in Pacific Islander populations-comprising only 0.002% of genome-wide association study (GWAS) participants as of 2024-remains unknown. We evaluated the transferability of multi-ancestry PGS for LDL cholesterol (LDL-C), HDL cholesterol (HDL-C), triglycerides (TGs), and total cholesterol (TC) in 4,342 Samoan adults across five cohorts spanning 1990-2010. PGSs from Graham et al. and Kanoni et al. multi-ancestry meta-analyses were harmonized with genome-wide imputed genotypes using a Samoan-specific reference panel, and performance was assessed via incremental R2 from linear mixed models with bootstrapped confidence intervals. HDL-C showed the highest performance (incremental R2 5.0%-15.0%), followed by TC (5.0%-10.7%), LDL-C (5.7%-8.6%), and TG (3.5%-7.0%). Critically, meaningful LDL-C performance was achieved only with the genome-wide PRS-CS score (99.6%-99.7% variant matching), while a curated pruning-and-thresholding score achieved ∼9% matching and near-zero performance. These findings establish systematic lipid PGS benchmarks in Samoans, demonstrating meaningful transferability when genome-wide variant coverage is ensured, and highlight variant harmonization as a critical precondition for PGS deployment in underrepresented populations.

Pacific Islanders

A module-based approach for post-omics, post-GWAS network-based gene classification.

MOTIVATION: Complex traits and diseases are highly polygenic and understanding the full set of genes involved is a central challenge in biomedicine. However, due to sample size limitations and noise (technical and biological), experimental approaches for disease-gene discovery such as transcriptomics and GWAS result in long, noisy, heterogeneous gene lists, which may be trimmed to a subset of likely relevant genes while leaving several false negatives. Computational gene classification approaches, especially those using genome-scale molecular interaction networks, are promising avenues for complementing such experimental findings by analytically expanding observed gene lists based on the functional relatedness between genes. We previously introduced the network-based gene classification approach, GenePlexus, which was rigorously benchmarked to show state-of-the-art performance, especially for predicting novel genes associated with biological processes and fine-grained phenotypes. Network-based gene classification performance,however, declines for diseases, especially when the inputs are omics and GWAS-based long gene lists. RESULTS: Here, we show that these disease gene lists span multiple biological processes spread across the molecular network, and we propose ModGenePlexus, a new network-based gene classification method that takes a two-stage approach. First, clustering and semi-supervised learning decomposes the input gene list into coherent, denoised network gene modules. Then, ModGenePlexus trains supervised (GenePlexus) classifiers for each module and aggregates predictions to return genome-wide rankings. We benchmarked ModGenePlexus across simulated data, transcriptomic signatures, and GWAS datasets (together spanning hundreds of diseases), showing improved recovery of known disease genes compared to GenePlexus. Beyond improved classification, the results of enrichment analysis of ModGenePlexus outputs are much more interpretable by virtue of revealing nuanced biological processes. Together, these results establish ModGenePlexus as a scalable, interpretable tool for gene classification of GWAS- and omics-derived gene lists across diverse biological contexts. AVAILABILITY AND IMPLEMENTATION: ModGenePlexus is freely available on GitHub at https://github.com/krishnanlab/ModGenePlexus, and the full source code and results supporting this study are available on Zenodo at https://zenodo.org/records/19857910.

Genome-Wide Association Study

scFANCL: Dual contrastive learning with false-negative correction at cell level for single-cell RNA-seq clustering.

BACKGROUND: Single-cell RNA sequencing (scRNA-seq) enables cellular characterization at single-cell resolution. However, its high dimensionality, sparsity, and noise make clustering challenging. Approaches utilizing contrastive learning and data augmentation have been introduced to improve representation quality for scRNA-seq clustering. In particular, dual contrastive frameworks combining instance- and cluster-level objectives can capture both cell-cell similarities and inter-cluster variations. However, existing dual contrastive frameworks focus primarily on discrete cluster boundaries, neglecting the biological continuity inherent in scRNA-seq data. METHODS: We propose scFANCL, a dual contrastive framework designed to capture biological continuity in scRNA data. Rather than treating all non-augmented samples as negatives, scFANCL applies a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, preserving continuous transcriptional relationships among them while maintaining inter-cluster separation. RESULTS: Extensive experiments across seven publicly available scRNA-seq datasets demonstrated that scFANCL achieves competitive clustering performance compared with existing baseline methods, consistently yielding high ARI and NMI scores across datasets of varying size and complexity. Ablation studies further confirmed the contribution of the false negative filtering component, showing measurable improvements over variants without filtering. Downstream analyses further suggest that the learned embeddings may reflect biologically meaningful transcriptional transitions, including continuous differentiation trajectories within related cell types. The source code is available at https://github.com/mjuailab/scFANCL . CONCLUSIONS: scFANCL addresses a key limitation of conventional contrastive learning by applying a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, thereby preserving biological continuity within cell types while maintaining inter-cluster separation. Evaluations across seven benchmark scRNA-seq datasets demonstrate competitive clustering performance, with learned embeddings capturing biologically meaningful transcriptional structure and characteristics of rare cell populations.

Clustering Algorithms

EPIPDLF: a pretrained deep learning framework for predicting enhancer-promoter interactions.

MOTIVATION: Enhancers and promoters, as regulatory DNA elements, play pivotal roles in gene expression, homeostasis, and disease development across various biological processes. With advancing research, it has been uncovered that distal enhancers may engage with nearby promoters to modulate the expression of target genes. This discovery holds significant implications for deepening our comprehension of various biological mechanisms. In recent years, numerous high-throughput wet-lab techniques have been created to detect possible interactions between enhancers and promoters. However, these experimental methods are often time-intensive and costly. RESULTS: To tackle this issue, we have created an innovative deep learning approach, EPIPDLF, which utilizes advanced deep learning techniques to predict EPIs based solely on genomic sequences in an interpretable manner. Comparative evaluations across six benchmark datasets demonstrate that EPIPDLF consistently exhibits superior performance in EPI prediction. Additionally, by incorporating interpretable analysis mechanisms, our model enables the elucidation of learned features, aiding in the identification and biological analysis of important sequences. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at: https://github.com/xzc196/EPIPDLF.

Deep Learning

FFC: a scalable FASTA compressor.

SUMMARY: FASTA is a widely used text-based format for storing nucleotide and protein sequences. The existing FASTA compressors usually focus on (slightly) improving the compression ratio, not on practical performance. We present FFC, a scalable FASTA compressor that achieves average compression speeds 4.7× and 11.4× higher than two high-performance compressors, zstd and NAF, respectively, across a benchmark set of seven single genomes. It also delivers average decompression speeds 3.5× and 2.7× higher than zstd and NAF, respectively. Although a chunk-based zstd variant with parallel decompression, pzstd, almost matches FFC speed, its compression ratio is on average by 23% worse than FFC's. For the experiment, a 14-core workstation and a RAM disk (to reduce the impact of I/O) were used. AVAILABILITY AND IMPLEMENTATION: FFC is freely available at github.com/kowallus/ffc and also as a Zenodo repository at 10.5281/zenodo.18892353, and the used datasets at 10.5281/zenodo.18873744.

Data Compression

Comprehensive evaluation of new sequencer T20 and well-established T7 with 507 human samples.

The DNBSEQ-T20×2 (T20) sequencer, developed by MGI Tech, enables cost-effective human whole-genome sequencing (WGS) at 30× coverage for less than $100 per genome. Here, we evaluate the sequencing performance and data quality of the T20 platform by benchmarking it against the established DNBSEQ-T7 (T7) sequencer using 507 samples derived from blood (N = 75), stool (N = 242), and saliva (N = 190). The T20 exhibited lower sequencing quality metrics compared with the T7, with Q20 scores of 95.76%-95.83% and Q30 scores of 87.25%-87.40%, compared with 97.81%-97.93% and 93.26%-93.60%, respectively, for T7 data. Quality differences were more evident toward the end of reads, and PCR-free libraries sequenced on the T20 showed similar reductions in quality scores. The median empirical base error rate estimated from 102 ZymoBIOMICS samples was 0.33%. The T20 demonstrated comparable coverage uniformity to the T7 and showed high concordance in microbiome composition analysis, with a median Bray-Curtis dissimilarity of 0.02. Variant calling performance was highly consistent between the two platforms. Among variants with non-missing genotype calls on both platforms, 94.92% of SNPs and 87.20% of InDels showed concordant genotypes between T20 and T7. Overall, the T20 delivers reliable sequencing accuracy and reproducibility for large-scale genomic and microbiome studies, providing a cost-effective alternative for high-throughput sequencing applications.

Metagenomics

CoMR: an integrative scoring pipeline for comprehensive mitochondrial proteome reconstruction across eukaryotes.

Mitochondrial proteome reconstruction from eukaryotic sequence data typically relies on prediction of mitochondrial targeting signals (MTSs). However, MTS predictors are primarily trained on model organisms and may perform poorly in phylogenetically divergent lineages or in organisms with atypical or reduced targeting sequences. Accurate reconstruction therefore requires integration of complementary sources of evidence beyond targeting prediction alone. We developed Comprehensive Mitochondrial Reconstructor (CoMR), an integrative workflow that combines targeting prediction, curated homology searches, large-scale similarity searches, and automated phylogenetic analysis within a unified scoring framework. Benchmarking on the model yeast Saccharomyces cerevisiae yielded strong discriminatory performance [receiver operating characteristic (ROC)-area under the curve (AUC) = 0.92], exceeding standalone prediction with TargetP2, a predictor of N-terminal targeting peptides (ROC-AUC = 0.72). In the divergent anaerobic protist Paratrimastix pyriformis, CoMR maintained robust performance (ROC-AUC = 0.86) validated with an experimental proteome despite extreme class imbalance, achieving a precision-recall AUC of 0.183 (~78-fold enrichment over random expectation and ~10-fold improvement over TargetP2). Ablation analyses demonstrate that predictive performance is robust to individual evidence-layer removal, while overlap analyses showed that homology-based searches recovered candidates missed by targeting predictors, particularly in P. pyriformis. Overall, CoMR improves mitochondrial proteome reconstruction over targeting prediction alone and provides a reproducible workflow for predicting mitochondrial and mitochondrion-related organelle protein repertoires across eukaryotes to aid investigations of organelle evolution and proteome reduction.

Proteome

Toward simple, rapid, and deep plant proteome analysis with an in-cell proteomics strategy.

While liquid chromatography-mass spectrometry (LCMS) has revolutionized plant proteomics over the past decade, plant sample preparation remains a major challenge due to rigid cell walls, abundant secondary metabolites, and wide dynamic range of protein abundance. These hurdles demand laborious tissue disruption, complex precipitation, and extensive cleanup prior to LCMS analysis, limiting the widespread adoption of proteomic technologies within the plant biology community. To overcome these barriers, we introduced an "in-cell proteomics" strategy that bypasses cell lysis and protein extraction by performing digestion directly inside methanol-fixed cells. We systematically benchmarked this strategy against conventional lysate-based workflows across 4 model plants (Arabidopsis thaliana, Nicotiana benthamiana, Zea mays, and Sorghum bicolor) and 3 tissue types (leaves, pollen, and seeds). Combined with minimal input material and single-shot LCMS, the in-cell approach consistently identified 9,000 to 12,000 proteins from leaves, 7,000 to 9,000 from pollen grains, and approximately 8,000 from seeds. Our comprehensive dataset demonstrates that this in-cell digestion approach substantially simplifies plant sample preparation while delivering proteomic performance equivalent to established workflows. Finally, to demonstrate the biological utility of this approach, we characterized the proteomes of N. benthamiana leaves infected with 2 fungal strains that exhibit different host specificities. Our in-depth proteomic data revealed distinct host response signatures differentiating the host-adapted Colletotrichum destructivum from the nonhost-adapted Colletotrichum sublineola strain. Overall, this study provides a simple, unbiased alternative for plant proteomic analysis that can be readily applied to tackle complex agricultural and physiological challenges in plant biology.

Proteomics

dcHiChIP: a comprehensive Nextflow-based pipeline for multiscale analysis of chromatin architecture from HiChIP data.

MOTIVATION: Despite the growing use of HiChIP to investigate protein-directed chromatin architecture, a comprehensive and reproducible pipeline for analysing these datasets-from raw reads to multiscale 3D genome features-remains lacking. Existing tools often focus on isolated components, such as loop calling or matrix generation, but fall short in integrating structural annotation, functional enrichment, and spatial modeling within a unified framework. To address this gap, we developed dcHiChIP, a modular, scalable Nextflow-based workflow that streamlines the analysis of HiChIP data, enabling both routine processing and in-depth exploration of chromatin organization and regulatory interactions. RESULTS: dcHiChIP enables robust and reproducible analysis of HiChIP datasets across multiple scales of chromatin architecture. It accepts raw sequencing data as input and generates high-quality loop calls, domain annotations, and 3D genome models. It also performs functional annotation and motif enrichment analyses. Applied to benchmark CTCF HiChIP datasets, dcHiChIP identifies major chromatin architectural features such as TADs/CCDs, A/B compartments, and chromatin stripes, and offers efficient, end-to-end execution with support for batch processing and workflow resumability. AVAILABILITY: dcHiChIP is publicly available on GitHub at https://github.com/SFGLab/dcHiChIP, with documentation at https://sfglab.github.io/dcHiChIP/. The software version used in this study is archived at Zenodo: https://doi.org/10.5281/zenodo.22030542.

Chromatin

Predicting host tropism in influenza a viruses: insights from multi-segment nucleotide signatures.

BACKGROUND: Influenza A virus (IAV) poses a significant public health threat due to its cross-species transmission and complex host adaptation mechanisms. This study integrated whole-genome data from avian, human, swine, and bovine IAV strains, using machine learning to predict viral host tropism based on nucleotide site features and to identify key sites driving host adaptation along with their synergistic effects. METHODS: A total of 64,000 IAV sequences from avian, human, swine, and bovine hosts were analyzed to build host-prediction models. A four-class classification framework (avian, human, swine, bovine) was constructed using nucleotide site features from all eight genomic segments (PB2, PB1, PA, HA, NP, NA, MP, NS). Eight machine learning algorithms (logistic regression, decision tree, random forest, SVM, KNN, gradient boosting, XGBoost, LightGBM) were benchmarked via 10-fold stratified cross-validation. Model performance was evaluated using accuracy, precision, recall, F1-score, AUPRC, and AUC. SHAP (SHapley Additive exPlanations) analysis prioritized critical nucleotide sites, while bivariate association tests identified synergistic/antagonistic interactions between sites. Nucleotide composition profiles were compared across host groups using hierarchical clustering and heatmap visualization. RESULTS: The XGBoost algorithm demonstrated the best and most stable performance, achieving an AUC value of over 0.95 in distinguishing human-derived sequences from non-human ones. SHAP analysis identified the top 20 critical nucleotide sites for each gene segment, such as sites 46 and 698 in the NS segment. Nucleotide composition analysis revealed high similarity between human and swine sequences in the HA and PB2 segments, and between avian and bovine sequences. The HA segment was particularly challenging in differentiating human from swine strains. Bivariate site association analysis uncovered significant synergistic or antagonistic effects between key sites within gene segments, forming complex networks. For instance, in the NS segment, a positive prediction contribution was observed when sites 371, 698, and 419 were all G. CONCLUSIONS: This study advances our mechanistic understanding of IAV host adaptation, identifies molecular determinants for zoonotic risk stratification, and establishes a scalable machine learning framework for predicting viral host tropism through nucleotide signature analysis, thereby enhancing surveillance strategies and informing preventive measures against emerging viral threats.

Influenza A virus

pLAST-a tool for rapid comparison and classification of bacterial plasmid sequences.

MOTIVATION: The increasing number of fully sequenced bacterial plasmids being annotated and catalogued has prompted the development of computational tools for comparing and classifying them. Existing approaches typically compare full-length DNA sequences (e.g. Mash, BLASTn, and ANI-based methods) or translated open reading frames (ORFs) (e.g. DIAMOND), with plasmid-level scores obtained by aggregating ORF-to-ORF similarities; however, they are either restricted to closely related plasmids or become computationally demanding in large-scale analyses. RESULTS: We describe pLAST (plasmid Language Analysis and Search Tool), a plasmid-search tool built using word2vec representations of protein-family content informed by local genomic context. Benchmarks indicate that pLAST outperforms nucleotide-based methods and performs comparably to DIAMOND in identifying functionally similar plasmids and compared with the widely used Mash, it achieves 26% and 24% improvements in detecting shared mating-pair formation system type and relaxase type, respectively. This performance scales to database searches across hundreds of thousands of sequences, as demonstrated using the precomputed PlasmidScope collection of ∼750 000 plasmids. Beyond global similarity, pLAST also returns per-ORF plasmid-plasmid alignments, enabling detection of shared functional modules. AVAILABILITY AND IMPLEMENTATION: pLAST is freely accessible as a web server at https://plast.lbs.cent.uw.edu.pl/ or https://plast.lbs.biol.uw.edu.pl/ and available as a Python module along with a precomputed database at https://github.com/labstructbioinf/pLAST for customized analysis.

Plasmids

Automated chromatin profiling with spa-ChIP-seq uncovers the impacts of condition variations.

Chromatin immunoprecipitation followed by sequencing (ChIP-seq) is widely used to study the genomic localization of DNA-associated proteins. However, conventional protocols include multiple manual steps that can introduce inconsistency and limit scalability, thereby restricting the inclusion of appropriate replicates and controls. Although the introduction of liquid handling platforms has improved reproducibility, most existing efforts have automated only a subset of the workflow, and extending automation to efficiently map non-histone proteins, such as chromatin regulators, remains challenging. Here, we present a fully automated implementation of our previously developed single-pot ChIP-seq protocol (Texari et al. 2021), named spa-ChIP-seq, which enables scalable processing of 8 to 96 ChIP-seq samples from crosslinked cells to sequencing-ready library in approximately three days with an estimated cost of $70 per sample. Benchmarking spa-ChIP-seq against manual ChIP-seq performed in parallel demonstrates comparable signal-to-noise ratio between the two workflows. Using spa-ChIP-seq, we systematically evaluate multiple parameters including shearing and crosslinking conditions, buffer compositions, and the ratio of antibody to cell-number. We find, for the first time to our knowledge, that weaker genomic localization signals are sensitive to changing the antibody to cell-number ratio, whereas the stronger signals remain unaffected. This finding underscores the importance of maintaining consistent antibody-to-cell-number ratio for comparative studies, such as treatment responses or chromatin-QTL mapping. The spa-ChIP-seq protocol is publicly available, including deck setups, operational parameters, and scripts. We envision that this robust, cost-efficient protocol will facilitate high-throughput, reproducible ChIP-seq analyses, supporting large-scale studies of antibody validation, compound screening, population genomics, and diagnostic frameworks.

Journal Article

Protein Language Model Decoys for Target Decoy Competition in Proteomics: Quality Assessment and Benchmarks.

Large-scale proteomics relies heavily on target-decoy competition for false discovery rate estimation in peptide identification, and the performance of this strategy depends strongly on the design of the decoy database. Classical generators such as reversal and shuffling remain widely used. Here, we introduce the first protein language model-based (PLM) decoy generation for peptide identification and benchmark it against classical strategies. We evaluate these approaches using three complementary quality-control layers: sequence-based separability, search-engine-agnostic spectral-space diagnostics, and end-to-end mass spectrometry benchmarks, including pipelines with rescoring. Across these analyses, PLM-based decoys are harder for sequence-only neural networks to distinguish than most classical generators, suggesting fewer obvious sequence-level artifacts. However, this signal is only weakly informative for search performance. Spectral diagnostics further show that short peptides occupy a particularly crowded target-decoy space and are therefore especially prone to local collisions across all generators. In full search pipelines, reverse decoys remain a strong baseline, and current PLM-based generators do not yet provide a clear overall advantage. We therefore view PLM-based decoys not as universal replacements for reverse decoys but as tunable tools for benchmarking, diagnostics, stress testing, and future adaptive decoy optimization, with increasing value as search models become more expressive.

Proteomics