PubMed HealthSearch

SEARCH · PubMed Health

Results for “High throughput workflow”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Systematic performance evaluation and application validation of an end-to-end NGS workstation.

Next-generation sequencing (NGS) library preparation is a core component of precision genomics, but it is commonly constrained by inefficiency, variability, and low throughput of manual protocols. To address these limitations, we developed and systematically evaluated a fully automated NGS workstations and further validated its performance across representative application scenarios. The automated system reduced total processing time from 8 to 10 to 4–6 h. At the same time, it maintained similar performance in pre-library metric, including DNA yield and fragment size, as well as post-capture sequencing metrics (Q30 > 90%, mapping rates > 95%, on-target rates 85–90%). The duplication rate was reduced to 5–8%, compared with 10–15% for manual methods, indicating increased library complexity. Bioinformatic evaluation of inter-species read mapping showed minimal cross-contamination, with a maximum contamination ratio of 0.0003%, indicating effective sample isolation in the automated workflow. High concordance in variant detection was observed between automated and manual workflows. Overall, this automated workstation provides a standardized and reproducible workflow that supports scalable precision genomics applications.

High-Throughput Nucleotide Sequencing

Accelerating natural product discovery, characterization and engineering by biofoundries.

Covering: From early developments to the presentNatural product (NP) discovery is increasingly constrained by low-throughput screening, repeated rediscovery, and challenges in scaling genome mining-guided validation workflows. This highlight examines how automated biofoundries are accelerating NP discovery, characterization, and engineering through integrated design-build-test-learn (DBTL) cycles. We discuss recent advances in phenotype-first and genome-first discovery strategies enabled by robotics, high-throughput pathway reconstitution, and automated screening platforms. We further highlight emerging technologies, including cell-free biosynthesis, automated culturomics, programmable chassis engineering, and AI-assisted workflow orchestration, that may enable increasingly autonomous biofoundries for scalable exploration of NP chemical space and therapeutic discovery.

Journal Article

Single-cell proteomics using mass spectrometry.

Over the past 2 to 3 years, mass-spectrometry-based single-cell proteomics (SCP) has experienced transformative improvements in microfluidic and robotic sample preparation, innovative MS1- and MS2-based multiplexing strategies, and specialized hardware (e.g., timsTOF Ultra 2, Astral), which have dramatically boosted sensitivity, throughput, and proteome coverage from picogram-level protein inputs. Concurrently, tailored computational workflows that encompass normalization, imputation, and no-code platforms have addressed pervasive missing data challenges and standardized analyses, collectively enabling high-throughput, reproducible profiling of cellular heterogeneity. This minireview summarizes the latest progress in SCP technology and software solutions, highlighting how the closer integration of analytical, computational, and experimental strategies will facilitate a deeper and broader coverage of single-cell proteomes.

Single-Cell Analysis

Hybridization capture increases on-target nanopore sequencing of plant RNA tobamovirus- derived cDNA libraries.

High-throughput sequencing (HTS) can support plant virus surveillance, but host nucleic acids often reduce on-target read recovery. We evaluated a targeted hybridization-capture workflow in which barcoded double-stranded cDNA (ds-cDNA) libraries generated from plant RNA extracts spiked with lyophilized tobamovirus-positive controls were enriched before Oxford Nanopore sequencing. Biotinylated probes targeted conserved regions of cucumber green mottle mosaic virus (CGMMV), species Tobamovirus viridimaculae; pepper mild mottle virus (PMMoV), species Tobamovirus capsici; and tobacco mosaic virus (TMV), species Tobamovirus tabaci. Across four pairs per virus, relative target-read abundance increased after capture from 0.76 ± 0.33% to 37.62 ± 15.72% for CGMMV, 8.16 ± 3.86% to 24.68 ± 12.34% for PMMoV, and 15.62 ± 10.40% to 36.83 ± 30.33% for TMV. Exact two-sided Wilcoxon signed-rank tests yielded P = 0.125 for each virus; with four nonzero differences in a common direction, this was the minimum attainable two-sided P value. Genome-coverage breadth was maintained. Retrospective duplex qPCR supported an increased virus-to-18S ratio for CGMMV, showed a variable PMMoV response, and showed a decreased virus-to-18S ratio for TMV because the 18S signal shifted earlier by as much as or more than the TMV signal. The findings provide proof-of-concept evidence for target-dependent library enrichment but do not establish analytical sensitivity, diagnostic performance, or field validity. Validation with naturally infected, low-titer, and mixed-infection samples and comparison with simpler targeted workflows are required.

biosecurity

NanoASV: a snakemake workflow for reproducible field-based Nanopore full-length 16S metabarcoding amplicon data analysis.

SUMMARY: NanoASV is a conda environment and snakemake-based workflow using state-of-the-art bioinformatics software to process full-length SSU rRNA (16S/18S) amplicons acquired with Oxford Nanopore Sequencing technology. Its strength lies in reproducibility, portability, and the possibility to run offline, allowing in-field analysis. It can be installed on the Nanopore MK1C sequencing device and process data locally. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are freely available at https://github.com/ImagoXV/NanoASV and Zenodo archive at https://doi.org/10.5281/zenodo.14730742.

Software

Colora: a Snakemake workflow for complete chromosome-scale de novo genome assembly.

MOTIVATION: De novo assembly creates reference genomes that underpin many modern biodiversity and conservation studies. Large numbers of new genomes are being assembled by labs around the world. To avoid duplication of efforts and variable data quality, we desire a best-practice assembly process, implemented as an automated portable workflow. RESULTS: Here, we present Colora, a Snakemake workflow that produces chromosome-scale de novo primary or phased genome assemblies complete with organelles using Pacific Biosciences HiFi, Hi-C, and optionally Oxford Nanopore Technologies reads as input. Colora is a user-friendly, versatile, and reproducible pipeline that is ready to use by researchers looking for an automated way to obtain high-quality de novo genome assemblies. AVAILABILITY AND IMPLEMENTATION: The source code of Colora is available on GitHub (https://github.com/LiaOb21/colora) and has been deposited in Zenodo under DOI https://doi.org/10.5281/zenodo.13321576. Colora is also available at the Snakemake Workflow Catalog (https://snakemake.github.io/snakemake-workflow-catalog/? usage=LiaOb21%2Fcolora).

Software

Proteoform profiling of endogenous single cells from rat hippocampus at scale.

We perform intact proteoform profiling of 10,809 endogenous single cells from the rat hippocampus using single-cell proteoform imaging mass spectrometry (scPiMS). scPiMS directly extracts whole proteins and demonstrates high throughput for MS-based single-cell proteomics compared with existing approaches. We develop an informatics workflow dedicated to this datatype and use it to assign neurons, astrocytes or microglia cell types according to their proteoform signatures.

Animals

Misdetection of frameshifts in SARS-CoV-2 genomes: need for additional harmonisation and efficient monitoring of data workflows.

Five years after the outbreak of the SARS-CoV-2 pandemic in 2020, diagnostic laboratories have moved from massive sequencing of thousands of samples to routine surveillance of SARS-CoV-2 cases, as with all other respiratory viruses. Surveillance remains of paramount importance to prevent a further SARS-CoV-2 surge, as the virus has been shown to mutate rapidly and can render available drugs and vaccines ineffective. During the pandemic, several bioinformatics pipelines and workflows have been developed to streamline analysis, shorten turnaround time and ensure reproducibility. As the number of samples decreases, laboratories are moving towards more flexible sequencing strategies and optimizing the cost per sample. However, workflow redesigns, even if individual steps have proven successful time and time again, can lead to challenges when changes in a bioinformatics pipeline are introduced (e.g. version updates, implementation of new features, etc.), a new combination of viral mutations emerge or a change in wet-lab procedures leads to unpredictable results. Here, we present a report of misidentified frameshift mutations in the consensus sequence of SARS-CoV-2, which led to an incorrect assumption of mutations in the spike and nucleocapsid viral proteins with the potential to affect PCR detection or even antigen testing. This investigation exemplifies the need for better awareness of the challenges that can occur even when using routinely applied protocols and analytical workflows and highlights the need for cooperation between experts of NGS, bioinformaticians and decision-makers towards more harmonized data workflows.

SARS-CoV-2

Reframing Proteomics Measurement: Super Mass Spectrometry Framework and the Role of Delayed Electrospray Ionization Technique.

Dynamic range, repeatability, and reproducibility remain the central limitations of data-independent acquisition (DIA) proteomics. Current workflows emphasize protein group identification counts and throughput, but these metrics mask the fundamental measurement challenge: generating a repeatable, reproducible, high-fidelity, and relatively complete digital representation of complex proteomes. In particular, plasma proteomics spans more than 10 orders of magnitude in protein abundance, far exceeding the capacity and dynamic range of any single mass spectrometer. Incremental advances have not closed this gap. In this Perspectives article, I introduce the Super Mass Spectrometry framework and then highlight the Delayed Electrospray Ionization (Delayed-ESI) technique, as a practical approach to address these limitations. By producing compositionally identical but temporally staggered ion beams, the Delayed-ESI technique enables deterministic remeasurement of the same analyte profile, supporting various novel strategies to improve analytical figures of merit. While recent implementations of the Delayed-ESI technique have emphasized throughput, I argue that the broader value of the Delayed-ESI technique lies in extending dynamic range and improving repeatability and reproducibility─objectives that should take precedence if proteomics is to evolve into a robust measurement science capable of supporting population-scale proteomics studies.

Proteomics

A statistical simulation model to guide the choices of analytical methods in arrayed CRISPR screen experiments.

An arrayed CRISPR screen is a high-throughput functional genomic screening method, which typically uses 384 well plates and has different gene knockouts in different wells. Despite various computational workflows, there is currently no systematic way to find what is a good workflow for arrayed CRISPR screening data analysis. To guide this choice, we developed a statistical simulation model that mimics the data generating process of arrayed CRISPR screening experiments. Our model is flexible and can simulate effects on phenotypic readouts of various experimental factors, such as the effect size of gene editing, as well as biological and technical variations. With two examples, we showed that the simulation model can assist making principled choice of normalization and hit calling method for the arrayed CRISPR data analysis. This simulation model is implemented in an R package and can be downloaded from Github.

CRISPR-Cas Systems

A Rapid Poly(ethylene glycol)-Assisted Magnetic Isolation Approach for High-Throughput Extracellular Vesicle Isolation and Subsequent Biomarker Analysis.

Extracellular vesicles (EVs) are crucial mediators of intercellular communication and have the potential to serve as biomarkers for disease diagnosis and therapeutic monitoring. However, most EV isolation methods often require large sample volumes and specialized instruments or involve trade-offs between purity, yield, cost, and scalability. We developed MagPEG, a workflow that combines poly(ethylene glycol) (PEG)-mediated EV aggregation with magnetic beads to provide a simple, reproducible alternative to ultracentrifugation, size-exclusion chromatography, and commercial precipitation kits. Our optimization experiments clarified the PEG concentration, ionic strength, and bead surface chemistry that collectively influence EV aggregation, capture efficiency, and contaminant coprecipitation, allowing us to define conditions that improve purity while maintaining high recovery. Compared with commonly used methods, MagPEG produced EVs with comparable size distribution, EV markers, and proteomic profiles while relying only on standard laboratory supplies. A key feature of the platform is that EVs and EV-associated DNA, RNA, and proteins can be sequentially extracted from the same bead-bound material, reducing sample loss and hands-on time and enabling multiomic analysis for limited clinical or small animal samples. MagPEG is compatible with downstream applications including proteomics, bead-based assays, and miRNA quantification. When applied to human serum, the method supported high-throughput EV proteomic profiling and enabled the identification of Alzheimer's disease-associated protein signatures, illustrating its utility for biomarker discovery. Overall, our results establish MagPEG as a powerful, rapid, scalable, and high-throughput solution for translational applications in biomarker discovery.

Polyethylene Glycols

Acoustic ejection mass spectrometry: the potential for personalized medicine.

INTRODUCTION: The emergence of personalized medicine (PM) has shifted the focus of healthcare from the traditional 'one-size-fits-all' approach to strategies tailored to individual patients, accounting for genetic, environmental, and lifestyle factors. Acoustic ejection mass spectrometry (AEMS) is a novel technology that offers a robust and scalable platform for high-throughput MS readout. AEMS achieves analytical speeds of one sample per second while maintaining high data quality, broad compound coverage, and minimal sample preparation, making it an invaluable tool for PM. AREAS COVERED: This article explores the potential of AEMS in critical PM applications, including therapeutic drug monitoring (TDM), proteomics, metabolomics, and mass spectrometry imaging. AEMS simplifies conventional workflows by minimizing sample preparation, enhancing automation compatibility, and enabling direct analysis of complex biological matrices. EXPERT OPINION: Integrating AEMS with orthogonal separation techniques such as differential mobility spectrometry (DMS) further addresses challenges in isomer discrimination, expanding the platform's analytical capabilities. Additionally, the development of high-throughput data processing tools could further enable AEMS to accelerate the development of personalized medicine.

Humans

Enzymes in high-throughput RNA sequencing: Applications and challenges.

High-throughput RNA sequencing provides genome-wide information on the dynamics of RNA in each cell and how the dynamics responds to environmental changes. Next-generation sequencing by the Illumina platform currently provides the highest information output as compared to other platforms. A key component of next generation sequencing of each RNA is the successful end-to-end reverse-transcription into a cDNA strand. This can be highly challenging given the propensity of each RNA to adopt ordered structures and to contain post-transcriptional modifications. While many reverse transcriptase (RT) enzymes have been developed over the years to maximize read-through of an RNA, their processivity and efficiency varies, raising the question of how to select the RT for the experiment at hand. Here, we use tRNA as a model for genome-wide sequencing, as tRNA has a stable secondary and tertiary structure and has a high density and wide variety of post-transcriptional modifications, presenting one of the most challenging problems of sequencing RNA. We compare the efficiency of end-to-end cDNA synthesis of tRNA among several recent RT enzymes and provide a general sequencing workflow that is applicable to most of these enzymes.

High-Throughput Nucleotide Sequencing

Evaluation of bone preparation approaches using length-based analysis and targeted sequencing for forensic human identification of historic skeletal remains.

Advances in DNA technology have significantly enhanced the forensic community's ability to develop genetic profiles from unidentified human skeletal remains. However, sampling requires mechanical grinding of hard tissues before DNA isolation. This processing can compromise genetic profiles, particularly in aged bones. We compared the industry-standard pulverization method with an alternative powder-free preparation involving prolonged demineralization and subsequent slicing of 19th-century cortical bone. Data from DNA quantification, STR genotyping, and targeted SNP sequencing were used to evaluate powdered samples versus demineralized slices from paired human bones. Average human DNA yields for pulverized samples and demineralized slices were 0.032&#x2009;ng and 0.692&#x2009;ng, respectively. Demineralized slices recovered more amplifiable DNA than traditional homogenization methods (p&#x2009;<&#x2009;0.05). No pulverized samples produced STR profiles, whereas demineralized slices from the same bone samples yielded partial profiles. Samples underwent DNA repair, library preparation, and hybridization capture using the FORensic Capture Enrichment (FORCE) panel. Applying low-coverage (1X) analysis of high-throughput sequencing (HTS) data, demineralized slices outperformed those prepared by traditional pulverization methods (p&#x2009;<&#x2009;0.05) and substantially increased the information recovered compared with conventional STR analysis methods. Based on HTS data from pulverized samples, DNA fragment length ranged from 27 to 95&#x2009;bp, and FORCE SNP recovery was 33.23%. In contrast, for demineralized slices, DNA fragment length ranged from 85 to 114&#x2009;bp, and FORCE SNP recovery was 83.24%. The required reagents and equipment are typically available in forensic labs, and the workflow outlined herein significantly increases the success of DNA recovery from challenging skeletal samples.

Humans

An Instrumental Optimization of a Label-Free Proteomic Method for Trace Protein Input.

Liquid chromatography-mass spectrometry (LC-MS)-based proteomics of trace-level samples, such as tens of cells or spatially resolved tissue regions, offers unique biological insights but is often constrained by the requirement for specialized, costly instrumentation. In this study, we developed a scalable workflow for the deep proteomic analysis of low- to ultralow-input samples by systematically optimizing a widely adopted Orbitrap and UHPLC platform to maximize sensitivity, precision, and throughput. This optimized workflow identified over 5600 proteins from 5 ng of peptides and 3400 proteins from 20 sorted cells, achieving a throughput of 30 analyses per day while maintaining deep proteome coverage and high quantitative reproducibility. Furthermore, by applying this method to spatially resolved proteomics, we identified over 6100 proteins from microscale regions of interest (ROIs) within a formalin-fixed, paraffin-embedded (FFPE) tissue. A data-driven normalization strategy was employed to correct for variable cellularity across tissue regions, effectively revealing intratumor heterogeneity and distinct molecular and functional signatures, including pathway activations not apparent in parallel spatial transcriptomic analysis. Ultimately, this accessible, high-performance method substantially lowers the instrumentation barrier for the deep proteomic profiling of trace-level biological samples.

Proteomics

ChemGenXplore: an interactive tool for exploring and analysing chemical genomic data.

MOTIVATION: Chemical genomics is a powerful high-throughput approach to systematically link phenotypes to genotypes. However, the vast datasets generated remain challenging to explore due to the lack of integrated, interactive tools for visualization and analysis. Existing workflows often require multiple independent software tools, limiting data accessibility and collaboration. Therefore, we created a user-friendly platform that enables efficient exploration and sharing of chemical genomics data. RESULTS: We developed ChemGenXplore, a web-based Shiny application designed to streamline the visualization and analysis of chemical genomic screens. It offers two primary functionalities: one for exploring pre-implemented datasets and another for analysing user-uploaded datasets. ChemGenXplore enables users to visualize phenotypic profiles, assess gene-gene and condition-condition correlations, perform GO and KEGG enrichment analysis, and generate customizable, interactive heatmaps. To further support collaborative research, ChemGenXplore also facilitates the comparative analysis of chemical genomic and other omics datasets. By consolidating these features into a single interactive and accessible tool, ChemGenXplore facilitates data sharing, enhances reproducibility, and promotes collaboration within the research community. AVAILABILITY AND IMPLEMENTATION: ChemGenXplore is freely accessible as a web application at https://chemgenxplore.kaust.edu.sa/. Source code and documentation, including instructions for local installation, are provided on GitHub (https://github.com/Hudaahmadd/ChemGenXplore). A Docker image is also available on DockerHub (https://hub.docker.com/r/hudaahmad/chemgenxplore) to ensure reproducibility and simplify installation.

Software

SynchroSep-MS: Parallel LC Separations for Multiplexed Proteomics.

Achieving high throughput remains a challenge in MS-based proteomics for large-scale applications. We introduce SynchroSep-MS, a novel method for parallelized, label-free proteome analysis that leverages the rapid acquisition speed of modern mass spectrometers. This approach employs multiple liquid chromatography columns, each with an independent sample, simultaneously introduced into a single mass spectrometer inlet. A precisely controlled retention time offset between sample injections creates distinct elution profiles, facilitating unambiguous analyte assignment. We modified the DIA-NN workflow to effectively process these unique parallelized data, accounting for retention time offsets. Using a dual-column setup with mouse brain peptides, SynchroSep-MS detected approximately 16,700 unique protein groups, nearly doubling the peptide information obtained from a conventional single proteome analysis. The method demonstrated excellent precision and reproducibility (median protein %RSDs less than 4%) and high quantitative linearity (median R2 greater than 0.96) with minimal matrix interference. SynchroSep-MS represents a new paradigm for data collection and the first example of label-free multiplexed proteome analysis via parallel LC separations, offering a direct strategy to accelerate throughput for demanding applications such as large-scale clinical cohorts and single-cell analyses without compromising peak capacity or causing ionization suppression.

Proteomics

DeepPlaque: a scalable multimodal platform for A&#x3b2; pathology and cell analysis in Alzheimer's disease.

Histological analysis is essential for understanding disease pathology and the microenvironment, particularly in Alzheimer's disease (AD), characterized by beta-amyloid (A&#x3b2;) plaques that exist as diffuse, fibrillar, and core species, with distinct toxicity levels. However, accurate classification of A&#x3b2; plaque types in postmortem brain tissues and profiling of surrounding cells present significant challenges. To address these challenges, we developed "DeepPlaque", an integrated system featuring "PlaqueNet", a deep learning model for automated classification of A&#x3b2; plaque species from diverse imaging platforms. DeepPlaque includes automated workflows for cellular phenotyping and proteomic profiling through targeted laser microdissection. PlaqueNet achieves expert-level accuracy (AUC&#x2009;>&#x2009;90%) in classifying the 3 major A&#x3b2; plaque species, supporting consistent and large-scale annotation. By integrating spatial cellular phenotyping with laser microdissection, DeepPlaque enables high-throughput proteomic analysis of A&#x3b2; plaque niches, revealing that microglia are more abundant around core and fibrillar A&#x3b2; plaques, with increased expression of apolipoprotein E and amyloid precursor protein in core A&#x3b2; plaques. This customizable platform enhances the molecular and cellular characterization of A&#x3b2; plaque-associated environments, providing critical insights into AD pathology.

Alzheimer Disease