PubMed HealthSearch

SEARCH · PubMed Health

Results for “benchmarking”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Benchmarking the OptiSpray-μPAC Workflow against a Traditional Nanospray Capillary Interface for Multiplexed Quantitative Proteomics.

Nanoflow liquid chromatography coupled with tandem mass spectrometry (LC-MS/MS) underpins modern quantitative proteomics, yet the column-to-mass spectrometer interface remains an important yet often underappreciated determinant of analytical depth, sensitivity, and reproducibility. Here, we benchmark an integrated workflow comprising the newly developed OptiSpray ion source and a micropillar array column (μPAC) cartridge against a conventional Nanospray Flex Source with an Accucore resin-packed capillary column. We performed a TMTpro 18-plex experiment across nine human cell lines on a FAIMS Pro-equipped Orbitrap Exploris 480. Following basic-pH reversed-phase fractionation, 12 fractions were analyzed on both workflow configurations under matched chromatographic gradient and acquisition conditions. Across both configurations, we quantified >9000 protein groups with highly comparable quantitative reproducibility and principal component clustering. Direct comparison of protein abundance ratios across cell lines showed agreement (Pearson R2 ≈ 0.7-0.8) without systematic bias. These results were achieved without workflow-specific optimization of the OptiSpray-μPAC platform, enabling direct transfer of established acquisition methods. Despite differences in column architecture, both configurations delivered comparable proteome coverage and quantitative fidelity. These findings establish the OptiSpray-μPAC workflow as a standardized alternative to conventional capillary-based interfaces, offering simplified operation while preserving quantitative performance.

Humans

Benchmarking methods for measuring biosynthetic gene cluster similarity and determination of gene cluster families.

MOTIVATION: Natural products are often produced by a set of biosynthetic enzymes that are encoded by genes clustered together in the producer's genome, referred to as a biosynthetic gene cluster (BGC). The ability to compare and cluster BGCs is essential for several applications, including predicting which bacteria will make a known product and assessing the potential diversity of natural products produced by a set of bacteria. There are multiple methods for comparing and clustering BGCs based on their similarity, but there has been a lack of investigation into how strongly BGC similarity relates to product structural similarity and how these methods perform relative to each other. RESULTS: Using publicly available databases, we developed a benchmark dataset to assess how well different BGC similarity metrics correlate with the structural similarity of their products and how well these methods cluster BGCs. We found that all methods showed moderate correlation between BGC and structural similarity, with correlations improving for more similar BGCs and varying significantly by BGC biosynthetic class. Analysis of outliers revealed some outliers were due to mistakes or omissions in public datasets, while others represented deviation between BGC similarity and product structural similarity. All methods generally performed better on clustering metrics, with BiG-SCAPE performing the best after errors in the public datasets had been corrected. AVAILABILITY AND IMPLEMENTATION: Scripts and data required to reproduce the results are available at https://github.com/aswalker-lab/BGC-clustering-benchmark and processed similarity, clusters, and scaffolds are also available at https://huggingface.co/datasets/allie-walker/BGC-clustering-benchmark. Code is also available at Zenodo: 10.5281/zenodo.17373546.

Multigene Family

Benchmarking with synthetic communities provides a baseline for virus-host inferences from Hi-C proximity linking.

Microbiomes influence diverse ecosystems, and viruses increasingly appear to impose key constraints. While viromics has expanded genomic catalogs, host identification for these viruses remains challenging due to the limitations in scaling cultivation-based approaches and the uncertain reliability and relative low resolution of in silico predictions - particularly for understudied viral taxa. Towards this, Hi-C proximity ligation uses sequenced, cross-linked virus and host genomic fragments to infer virus-host linkages and has now been applied in at least 10 studies. However, its accuracy remains unknown. Here we assess Hi-C performance in recovering virus-host interactions using synthetic communities (SynComs) composed of four marine bacterial strains and nine phages with known interactions and then apply optimized bioinformatic protocols to natural soil samples. In SynComs, standard Hi-C sample preparations and analyses showed poor normalized contact score performance (26% specificity, 100% sensitivity, incorrect matches up to class level) that could be dramatically improved by Z-score filtering (Z ≥ 0.5, 99% specificity), though at reduced sensitivity (62% down from 100%). Detection limits were established as reproducibility was poor below minimal phage abundances of 105 PFU/mL. Applying optimized bioinformatic protocols to natural soil samples, we compared virus-host linkages inferred from proximity-ligated Hi-C sequencing with predictions generated by in silico homology-based and machine learning-based bioinformatic approaches. Prior to Z-score thresholding, agreement was relatively high at the phylum to family levels (72%), but not at the genus (43%) or species (15%) levels. Z-score thresholding reduced sensitivity (only 34% of predictions were retained), with only modest improvements in congruence with bioinformatic methods (48% or 18% at genus or species levels, respectively). Regardless, this led to 79 genus-level-congruent virus-host linkages and 293 new ones revealed by Hi-C alone, i.e., providing many new virus-host interactions to explore in already well-studied climate-critical soils. Overall, these findings provide empirical benchmarks and methodological guidelines to improve the accuracy and reliability of Hi-C for virus-host linkage studies in complex microbial communities.

Benchmarking

Benchmark of biomarker identification and prognostic modeling methods on diverse censored data.

The practices of identifying biomarkers and developing prognostic models using genomic data has become increasingly prevalent. Such data often features characteristics that make these practices difficult, namely high dimensionality, correlations between predictors, and sparsity. Many modern methods have been developed to address these problematic characteristics while performing feature selection and prognostic modeling, but a large-scale comparison of their performances in these tasks on diverse right-censored time to event data (aka survival time data) is much needed. We have compiled many existing methods, including some machine learning methods, several which have performed well in previous benchmarks, primarily for comparison in regards to variable selection capability, and secondarily for survival time prediction on many synthetic datasets with varying levels of sparsity, correlation between predictors, and signal strength of informative predictors. For illustration, we have also performed multiple analyses on a publicly available and widely used cancer cohort from The Cancer Genome Atlas using these methods. We evaluated the methods through extensive simulation studies in terms of the false discovery rate, F1-score, concordance index, Brier score, root mean square error, and computation time. Of the methods compared, CoxBoost and the Adaptive LASSO performed well in all metrics, and the LASSO and elastic net excelled when evaluating concordance index and F1-score. The Benjamini-Hoschberg and q-value procedures showed volatile performances in controlling the false discovery rate. Some methods' performances were greatly affected by differences in the data characteristics. With our extensive numerical study, we have identified the best performing methods for a plethora of data characteristics using informative metrics. This will help cancer researchers in choosing the best approach for their needs when working with genomic data.

Humans

Spinal meningiomas: histopathological grading using a benchmark radiomics model with notes on disease control.

OBJECTIVE: Spinal meningiomas (SMs) are common primary spinal tumors for which surgery is considered the first-line treatment when safe and feasible. The ability to extrapolate the tumor grade from preoperative imaging may significantly inform early patient expectation-setting regarding recurrence. Building on radiomics studies in cranial meningiomas, the authors aimed to construct a benchmark radiomics model to preoperatively identify the histological grade of SMs. METHODS: Institutional surgical records from May 2012 to November 2025 were queried for pathology-confirmed meningiomas below the foramen magnum, with preoperative contrast-enhanced imaging available for segmentation. SMs were classified as low-grade (WHO grade 1) and high-grade (WHO grade 2 tumors and grade 1 tumors with atypia). Tumors were manually segmented, and features were extracted using the PyRadiomics software package. An ensemble model of k-nearest neighbors, random forest, and support vector machine classifiers was trained using nested cross-validation on a subset of 10 features to differentiate tumor grades. Clinical data for the cohort were also extracted, and disease control in an adjunctive clinical series was assessed. RESULTS: Seventy-four patients were included in radiomics analysis, with an area under the receiver operating characteristic curve of 0.879 and a mean F1 score of 0.748. The model's top 5 features were all texture features that differed significantly (p < 0.05) across low- and high-grade SMs. These included measures of tumor textural and contrast-enhancement heterogeneity, with overlap with features reported in radiomics models for histological grading of intracranial meningiomas. Fifty-five patients with a median radiographic follow-up of 22.2 (range 1.9-86.4) months remained for clinical analysis after exclusion of patients with less than 1 month of follow-up and syndromic meningiomas. Four recurrences occurred at a median of 20.8 (range 1.8-41.8) months. High-grade tumor pathology did not significantly impact progression-free survival (p = 0.682, log-rank test; Cox regression high vs low grade hazard ratio [HR] 0.62, 95% CI 0.06-6.11, p = 0.685). Subtotal resection was associated with poorer progression-free survival than gross-total resection (p = 0.004, log-rank test; Cox regression subtotal vs gross-total resection HR 10.62, 95% CI 1.46-77.05, p = 0.019). These findings remain contextualized within a relatively limited follow-up window and small recurrence event count, suggesting a need to characterize the interplay between tumor grade and extent of resection as drivers of local disease control in SMs. CONCLUSIONS: A preoperative radiomics model can stratify high-grade SMs using open-source tools applied to single-institution data.

Humans

Rallpacks: a set of benchmarks for neuronal simulators.

The field of computational neurobiology has advanced to the point where there are several general-purpose simulators to choose from. These cater to various niches in the world of realistic neuronal models, which range from the molecular level to descriptions of entire sensory modalities. In addition, there are numerous custom-designed simulations, adaptations of electrical circuit simulators, and other specific implementations of neurobiological models. As a first step towards evaluating this disparate set of simulators and simulations, and towards establishing standards for comparisons of speed and accuracy, we describe a set of benchmarks. These have been given the name 'Rallpacks' in honor of Wilfrid Rall, who pioneered the study of neuronal systems through analytical and numerical techniques.

Algorithms

Genome-sequencing-based benchmarking of antimicrobial resistance, treatment outcomes and healthcare transmission events for Clostridioides difficile infection in Australian hospitals.

BACKGROUND: Clostridioides difficile infection (CDI) remains a priority for infection prevention and control in health care, particularly with the emergence of hypervirulent strains and antimicrobial resistance (AMR). AIM: To characterize the genomic epidemiology and AMR profiles of culture-confirmed CDI cases within tertiary hospitals in Australia. METHODS: A total of 155 C. difficile isolates from 142 patients with CDI diagnosed in four hospitals between 2023 and 2025 were studied. Data collected included patient demographics, severity of infection, antibiotic treatment and clinical outcomes at 8 weeks. Phenotypic susceptibility to vancomycin, fidaxomicin, metronidazole, moxifloxacin, meropenem, tetracycline and rifaximin were determined by agar dilution. Isolates underwent whole-genome sequencing (WGS) for genotyping and resistome assessment. FINDINGS: WGS differentiated 39 distinct sequence types among CDI isolates across different healthcare services. In total, 100 isolates were singletons and 55 (35% clustering rate) isolates were considered to be genomically related (difference of two or fewer single-nucleotide polymorphisms). Of these, 12 patients (8.5%) with close hospital contact formed six epidemiologically linked clusters. Phenotypic susceptibility results were obtained for 134 (86.4%) CDI isolates. There was no phenotypic resistance to vancomycin [minimum inhibitory concentration required to inhibit the growth of 90% of isolates (MIC90) 1 mg/L], metronidazole (MIC90 0.5 mg/L) or fidaxomicin (MIC90 0.5 mg/L). There was no association in the study cohort between the presence of resistance genes or reduced phenotypic susceptibility and CDI recurrence. CONCLUSION: Genomic analysis of C. difficile isolates did not identify any outbreaks or an association between the sequence type or presence of a resistance gene and clinical outcomes. High-resolution characterization and identification of antibiotic resistance, CDI clinical relapse and recent transmission offered by genome sequencing can provide important benchmarks for hospital infection control.

Antibiotic resistance

Benchmark for Quantitative Global and Redox Proteomics Analysis by Combining Protein-Aggregation Capture and Data Independent Acquisition.

Oxidative damage plays a critical role in various diseases including cardiovascular and neurological disorders. Thiol redox reactions, acting as oxidative stress sensors, influence protein structure and function. Redox proteomics, based on the differential alkylation of cysteine sites followed by mass spectrometry, enables the comprehensive analysis of thiol redox status in cells and tissues. However, these approaches require extensive sample manipulation and are not compatible with data-independent acquisition techniques. Here, we introduce PACREDOX, an innovative strategy based on protein aggregation capture (PAC), and demonstrate its compatibility with library-free DIA. Compared with traditional methods such as FASILOX, PACREDOX reduces preparation time and costs while maintaining thiol and proteome coverage. To enable library-free DIA, we corrected in silico spectral libraries in DIA-NN using experimental retention time data from methylthiolated-Cys peptides. PACREDOX with DIA was benchmarked against FASILOX in a myocardial infarction model, yielding the same biological insights, while enhancing peptide and protein coverage. Our results underscore the potential and efficiency of this methodology for studying oxidative damage. Overall, PACREDOX offers an automatable, high-throughput, and cost-effective strategy for redox proteomics.

Proteomics

Benchmark test of transport calculations of gold and nickel activation with implications for neutron kerma at Hiroshima.

A benchmark test of the Monte Carlo neutron and photon transport code system (MCNP) was performed using a 252Cf fission neutron source to validate the use of the code for the energy spectrum analyses of Hiroshima atomic bomb neutrons. Nuclear data libraries used in the Monte Carlo neutron and photon transport code calculation were ENDF/B-III, ENDF/B-IV, LASL-SUB, and ENDL-73. The neutron moderators used were granite (the main component of which is SiO2, with a small fraction of hydrogen), Newlight [polyethylene with 3.7% boron (natural)], ammonium chloride (NH4Cl), and water (H2O). Each moderator was 65 cm thick. The neutron detectors were gold and nickel foils, which were used to detect thermal and epithermal neutrons (4.9 eV) and fast neutrons (> 0.5 MeV), respectively. Measured activity data from neutron-irradiated gold and nickel foils in these moderators decreased to about 1/1,000th or 1/10,000th, which correspond to about 1,500 m ground distance from the hypocenter in Hiroshima. For both gold and nickel detectors, the measured activities and the calculated values agreed within 10%. The slopes of the depth-yield relations in each moderator, except granite, were similar for neutrons detected by the gold and nickel foils. From the results of these studies, the Monte Carlo neutron and photon transport code was verified to be accurate enough for use with the elements hydrogen, carbon, nitrogen, oxygen, silicon, chlorine, and cadmium, and for the incident 252Cf fission spectrum neutrons.

Californium

Synthetic community Hi-C benchmarking provides a baseline for virus-host inferences.

Microbiomes influence diverse ecosystems, and viruses increasingly appear to impose key constraints. While viromics has expanded genomic catalogs, host identification for these viruses remains challenging due to the limitations in scaling cultivation-based approaches and the uncertain reliability and relative low resolution of in silico predictions - particularly for understudied viral taxa. Towards this, Hi-C proximity ligation uses sequenced, cross-linked virus and host genomic fragments to infer virus-host linkages and has now been applied in at least ten studies. However, its accuracy remains unknown. Here we assess Hi-C performance in recovering virus-host interactions using synthetic communities (SynComs) composed of four marine bacterial strains and nine phages with known interactions and then apply optimized bioinformatic protocols to natural soil samples. In SynComs, standard Hi-C sample preparations and analyses showed poor normalized contact score performance (26% specificity, 100% sensitivity, incorrect matches up to class level) that could be dramatically improved by Z-score filtering (Z &#x2265; 0.5, 99% specificity), though at reduced sensitivity (62% down from 100%). Detection limits were established as reproducibility was poor below minimal phage abundances of 105 PFU/mL. Applying optimized bioinformatic protocols to natural soil samples, we compared virus-host linkages inferred from proximity-ligated Hi-C sequencing with predictions generated by in silico homology-based and machine learning-based bioinformatic approaches. Prior to Z-score thresholding, agreement was relatively high at the phylum to family levels (72%), but not at the genus (43%) or species (15%) levels. Z-score thresholding reduced sensitivity (only 34% of predictions were retained), with only modest improvements in congruence with bioinformatic methods (48% or 18% at genus or species levels, respectively). Regardless, this led to 79 genus-level-congruent virus-host linkages and 293 new ones revealed by Hi-C alone - i.e., providing many new virus-host interactions to explore in already well-studied climate-critical soils. Overall, these findings provide empirical benchmarks and methodological guidelines to improve the accuracy and reliability of Hi-C for virus-host linkage studies in complex microbial communities.

Genomics

Evaluating Language Models for Biomedical Fact-Checking: A Benchmark Dataset for Cancer Variant Interpretation Verification.

Accurate interpretation of genomic variants is critical for precision oncology but remains slow and dependent on specialized expertise. Public knowledgebases such as the Clinical Interpretation of Variants in Cancer (CIViC) help by curating literature-backed variant interpretations in a structured form, yet verification and review have become major bottlenecks. To address this, we developed CIViC-Fact, a benchmark dataset and pipeline for testing automated systems that verify the accuracy of cancer variant claims. CIViC-Fact links structured claims to sentence-level supporting or refuting evidence from full-text articles, and includes expert annotations and explanations. We evaluated multiple language models. Proprietary models performed well without training, but a smaller open-source model, fine-tuned on CIViC-Fact, achieved the highest accuracy (89%). Applying our fact-checking pipeline to real CIViC entries showed that reviewing less than 20% of content, focusing on flagged entries, would be sufficient to catch over half of all errors. This AI-assisted triage greatly accelerates the review process without replacing or reducing expert insight, ensuring that existing careful oversight remains in place while curators can work more efficiently. CIViC-Fact provides a realistic, high-consequence framework for biomedical fact-checking and a path toward more rigorous and efficient knowledgebase curation.

Journal Article

Transcriptional benchmark dose modeling of ultraviolet radiation-induced genomic activation in mouse skin.

The in&#xa0;vivo transcriptional response of mouse skin to ultraviolet radiation (UV-R) exposure reveals key genomic alterations associated with UV-R-induced damage but it does not provide precise dose thresholds for these effects. These initial findings provided the impetus to advance dose-response characterization by integrating benchmark dose (BMD) modeling with transcriptomic data, aiming to identify biologically relevant points of departure for gene and pathway activation. To accomplish this, mice were exposed to five erythemally weighted UV-R doses (0-40&#x2009;mJ/cm2) emitted from a UV-emitting tanning device, across six post-exposure timepoints (0-96&#x2009;h). Four analytical methods were used to estimate BMDs, with the lowest consistent response dose (LCRD) approach yielding the most sensitive estimates (1.21-3.44&#x2009;mJ/cm2). Transcriptomic responses revealed activation of shared pathways related to DNA damage and cancer, oxidative stress and metabolism, inflammation and immunity, and hormonal disruption. Notably, the majority of LCRD BMD estimates (1.21-3.44&#x2009;mJ/cm2) were lower than the International Electrotechnical Commission standard actinic exposure limit (3&#x2009;mJ/cm2 (erythemally weighted)) for broadband UV-R (200-400&#x2009;nm) for unprotected skin and the eye for an 8&#x2009;h period. These findings suggest that transcriptomic BMD modeling can detect early biological responses to UV-R at doses lower than current exposure limits.

Animals

Benchmarking: a tool for excellence in palliative care.

Quietly and without fanfare, total quality management (TQM) is being implemented in a branch of health care where quality of care has particular impact on the patient's comfort and well-being. Some palliative care providers, dedicated to improving the quality of life for the dying, have fulfilled all the criteria to be contenders for prestigious quality honors like the Baldrige Award in the United States and the Canada Award for Excellence. Their secret is simple: the patient defines quality, and the palliative care team acts on that definition. Benchmarking, a TQM tool, allows institutions and organizations to benefit from sharing their best processes, and keeps the TQM continuous improvement cycle on track.

Home Care Services

Large-scale benchmarking of prokaryotic annotation tools across thousands of species.

BACKGROUND: Genome annotation is an important step in deriving functional meaning from prokaryotic sequencing data, yet systematic evaluations guiding tool selection are lacking. We present the first large-scale investigation of four prominent open-source annotation tools (Prokka, Bakta, EggNOG-mapper, and PGAP) across 156,033 diverse genomes. This includes Escherichia coli strains for baseline performance, thousands of archaea and bacteria genomes, as well as frameshifted and metagenome-assembled genomes. RESULTS: Bakta excels in annotating high-quality bacterial genomes, while PGAP was better for archaeal genomes and challenging bacterial assemblies, including metagenome-assembled, fragmented, or contaminated samples. For Gene Ontology annotation, PGAP consistently provides broader term coverage, whereas EggNOG-mapper offers more terms per feature. CONCLUSIONS: Our findings highlight tool-specific strengths crucial for selecting optimal solutions based on genome quality, taxonomy, and origin (e.g. MAGs). This study provides an evidence-based guide for users and informs future tool development.

Molecular Sequence Annotation

Response to: "best practices when benchmarking CATCH for the design of genome enrichment probes".

We clarify the design principles and evaluation choices underlying Syotti, a robust and scalable probe-design tool developed to support large, heterogeneous bacterial datasets with minimal parameter tuning. We highlight Syotti's ability to perform simultaneous large-scale designs and its effectiveness as a reliable alternative when existing tools such as CATCH are not well suited to the problem setting.

Genomics

Benchmarking DNA extraction protocols across use cases for culture-independent Nanopore metagenomics.

Oxford Nanopore Technologies (ONT) sequencing offers several advantages for metagenomics, including long reads, rapid turnaround, low upfront cost, scalability and portability. However, for ONT metagenomics, DNA yield, quality and integrity are important considerations when selecting an extraction method. Many metagenomic extraction methods use harsh lysis conditions to extract a wide range of species and provide an accurate community composition, but these conditions can compromise DNA fragment length. Therefore, extraction methods for ONT metagenomics must balance DNA shearing and recovery with representative community lysis. We systematically evaluated DNA extraction methods for ONT metagenomic sequencing using a use case-oriented framework. Among nearly 50 extraction methods screened, 7 were selected for detailed comparison based on suitability for metagenomics, variation in methodology, availability, cost and processing time: Norgen BioTek Corp's Stool DNA Isolation (NG), Zymo Research's ZymoBIOMICS Quick-DNA HMW MagBead (ZMG), Qiagen's DNeasy Blood and Tissue (QBT), Macherey-Nagel's NucleoMag DNA Microbiome (MN), Zymo Research's ZymoBIOMICS DNA Mini Prep (ZMI), Qiagen's DNeasy PowerSoil/QIAamp PowerFecal Pro (PS) and Qiagen's QIAamp Fast DNA Stool Mini (QIA). Methods were tested using Zymo Research's ZymoBIOMICS Microbial Community Standard (MCS), a matrix-free mock community with known composition. DNA extracts were sequenced on an ONT PromethION using the Rapid Barcoding Kit, except QIA due to insufficient DNA yield. Metrics for the method, DNA extracts, sequencing and genomes were evaluated, revealing trade-offs between methods. The two magnetic bead methods, MN and ZMG, produced the highest mean read length N50 values (13.9 and 16.5&#x2009;kb, respectively) but showed apparent community compositions skewed towards Gram-negative bacteria. In contrast, ZMI and PS maintained a community composition close to expected, with reduced mean read length N50 values (4.5 vs. 7.5&#x2009;kb). Performance across various metrics is presented in the context of the following use cases: maximizing genome coverage and assembly completeness, preserving composition accuracy, targeting specific species and limiting required resources (equipment, time or budget). The metrics and use case considerations presented offer practical guidance for informed selection of DNA extraction methods for ONT metagenomics. For accurate community composition, ZMI or PS are recommended, while PS and ZMG perform best at maximizing genome coverage and assembly completeness. NG and QBT may be the most economical options, though performance trade-offs were observed. Finally, PS may be the preferred method for time-sensitive diagnostic or field applications.

Metagenomics

A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings.

In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime.

Humans

Benchmarking computational decontamination of ambient RNA.

Gene expression profiling of single cells using single-cell and single-nucleus RNA sequencing (sxRNA-seq) enables researchers to characterize cellular heterogeneity and unraveling complex biological processes at unprecedented resolution. However, sxRNA-seq faces challenges due to the presence of ambient RNA, extraneous RNA molecules not originating from the cells of interest. Sample preparation is a major source of ambient RNA, where harsh conditions can lead to cell lysis and the release of intracellular RNA. This inescapable inclusion of ambient RNA can cause erroneous results and hinder downstream analyses. To address this issue, various methodologies have been developed to identify, quantify, and remove ambient RNA. Here, we rigorously evaluate 7 state-of-the-art methodologies for ambient RNA removal using simulated datasets, species-mixing experiments of varying complexities, and genotype-mixing experiments. We find that no single method performs the best across all datasets and metrics, but CellBender, DecontX and SoupX generally perform well.

ambient RNA