PubMed HealthSearch

SEARCH · PubMed Health

Results for “workflow”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A portable recalibration workflow for reference-based variant calling in non-human genomes.

A key computational step in reference-based variant calling is distinguishing true genetic variants from sequencing errors. Advanced tools and workflows have been developed to handle this by computational modelling of technical errors from the sequencing machines. However, these recalibration workflows have largely been evaluated for human data only and its exact applicability for non-human data remains unknown. Here, we conducted a systematic evaluation of variant calling on human, rice, sheep, and chickpea data, and found that existing workflows introduce unexpected statistical bias, thus leading to suboptimal variant calls for non-human data. To address this problem, we present simple guidelines for constructing a "pseudo-"database (pseudoDB) of genetic variants as a scalable and portable solution for recalibration and variant calling. With human data, our pseudoDB-based workflow performs comparably to existing dbSNP-based GATK3 workflows and those using DeepVariant, Strelka2, and FreeBayes. We extend this to other non-human genomes, namely cattle, brown bear, swan goose, African oil palm, Komodo dragon, and stevia, altogether resulting in the identification of up to 242.0% unique genetic variants. The majority of newly identified variants are within the non-coding regions, hinting at the rich diversity of genome regulation in the non-human population. Our pseudoDB-based workflow is agnostic to reference genomes and modular for easy integration with other computational workflows for human and non-human resequencing data.

Humans

A streamlined workflow for high throughput metaproteomic analysis of the rumen microbiome.

Metaproteomics can provide direct functional insights into complex microbial communities, yet its application in rumen research remains limited due to labor-intensive and low-throughput sample preparation workflows before the MS analysis. This work aimed to develop and characterize a streamlined, high throughput metaproteomic workflow optimized for rumen samples. Key steps, including microbial cell extraction, cell lysis, protein digestion, and LC-MS/MS acquisition, were systematically assessed and optimized to reduce hands-on time while maintaining deep proteome coverage. The optimized workflow integrates a minimized cell extraction protocol using 0.5 g starting material and in-solution tryptic digestion. Application of the final workflow to 72 samples from in vitro fermentation revealed that biological variability between inocula dominated technical variability, which remained moderate (median CV of 21-24% across batches). Overall, the optimized workflow supports robust taxonomic and functional characterization of the rumen microbiome with improved scalability. These advances provide a foundation for applying metaproteomics to larger experimental designs, including nutritional trials and cohort studies, thereby enabling broader functional interrogation of rumen microbial ecosystems. SIGNIFICANCE: This study addresses current limitations in the application of metaproteomics to rumen microbiome research by developing a streamlined and scalable sample preparation workflow. By optimizing key steps and reducing sample input while maintaining reproducibility and proteome coverage, this work enables more efficient processing of larger sample sets. These advances support the broader use of metaproteomics in rumen studies and facilitate functional investigations relevant to animal nutrition and sustainable livestock production.

Animals

Comparative performance of portable DNA extraction protocols and bioinformatics workflows for rapid detection of gram-negative bacteria and antimicrobial resistance using Oxford Nanopore sequencing.

Oxford Nanopore Technology (ONT) enables rapid, portable pathogen identification and antimicrobial resistance (AMR) detection, but the reliability of downstream genomic analyses is highly dependent on DNA extraction quality, particularly in resource-limited settings. This study comparatively evaluated four portable bacterial DNA extraction protocols derived from three commercial kits to determine their impact on nanopore sequencing performance, bioinformatics workflow completion, and field deployability. Six gram-negative bacterial isolates (Escherichia coli, n = 4; Pseudomonas sp., n = 1; and Salmonella sp., n = 1) were processed using four extraction protocols: SwiftX DNA, SwiftX DNA with proteinase K (ProtK), SwiftX ParaBact, and NucleoSpin Microbial. Twenty-four resulting DNA extracts were sequenced on a single multiplexed MinION R10.4.1 flow cell. Sequencing data were analyzed using validated Galaxy-based generic and species-specific pipelines. Workflow completion was defined as successful progression through quality control, assembly, virulence, plasmid, and AMR detection modules. DNA purity varied substantially by extraction protocol and was strongly associated with successful workflow completion (Kruskal-Wallis, P = 0.0006). Accordingly, NucleoSpin Microbial achieved 100% workflow completion, and SwiftX ParaBact achieved 83%, while both SwiftX DNA-based protocols failed to complete full workflows. Importantly, key AMR genes required to classify isolates as multidrug-resistant were consistently detected using both NucleoSpin Microbial and SwiftX ParaBact extractions. However, NucleoSpin Microbial assemblies showed significantly higher contiguity and enabled a broader, more complete detection of virulence factors, pathogenicity islands, plasmid replicons, and accessory AMR genes, reflecting enhanced genomic resolution.IMPORTANCERapid whole-genome sequencing is increasingly used to detect antimicrobial resistance and guide public health responses, but its reliability depends strongly on how bacterial DNA is extracted. In this study, we have shown that DNA extraction method choice has a major impact on Oxford Nanopore sequencing performance across clinically relevant gram-negative bacteria. While silica column-based extraction maximized genomic completeness and analytical depth, paramagnetic bead-based reverse purification offered superior portability with sufficient resolution for frontline AMR surveillance. These findings highlight a practical trade-off between field deployability and high-resolution genomic characterization in low-resource settings.

DNA extraction

A Simplified Workflow for the Prediction of Putative Viral Reads Using NIPT Data.

OBJECTIVE: Non-invasive prenatal testing (NIPT) identifies fetal chromosomal abnormalities by sequencing cell-free fetal DNA (cffDNA). Recent studies suggest the prediction of viral sequences from NIPT data, but current methods lack cost-effectiveness for routine use. This study develops a straightforward workflow to investigate potential viral signatures in pregnant women using NIPT data from 888 Iranian participants. METHOD: Two bioinformatic workflows were compared for predicting viral reads: the traditional method involved mapping reads to the human genome, followed by mapping unmapped reads to viral references, and a direct mapping approach to viral genomes, as proposed in this research. RESULTS: While maintaining reproducibility comparable to the conventional method, the proposed workflow minimizes computational complexity and time usage for data processing. Ultimately, this analysis suggested viral DNA in 24.2% of samples, encompassing 29 distinct species, implying the diversity of the maternal virome. CONCLUSION: This study presents a computationally efficient workflow for the in silico prediction of viral-like sequences from routine NIPT data. Further experimental validation is essential to verify the presence, viability, or clinical relevance of these sequences.

Humans

Experimental workflows for the accurate identification of mitochondrial redox events.

The study of redox biology has been growing constantly since the last decades. Over these years, redox processes have been linked to an extraordinarily wide range of physiological and pathological events, becoming recognized as central mechanisms underlying many of them. In this context, it becomes essential to understand the advantages and limitations of the tools under use, to recognize the specific controls required for each measurement and to accurately distinguish between distinct redox mechanisms. So far, multiple and excellent reviews have dealt with either the tools, the protocols or the mechanisms involved in reactive oxygen species (ROS) production and quenching, a.k.a. redox events. However, a review outlining the workflows to appropriately detect them is still lacking. We define workflow as the combination of tools, methods and mechanistic knowledge that allow the definition of a specific redox event. In this review, we aim to provide an optimal workflow for the research on mitochondrial redox events. To this end, we first summarize the molecular tools available to measure and quench ROS. We then explain the mechanisms of ROS production and scavenging in several of the cellular compartments, with special focus on mitochondria, as well as their implication in physiology and disease. Finally, we use the knowledge in all sections to build a recommended experimental workflow, illustrated by several cases of study. This review will enable the reader to understand how specific mitochondrial redox events can be accurately measured, considering all technical, methodological and mechanistical variables and limitations required for their reliable detection and interpretation.

(Patho)physiology

A novel reusable transcriptome-wide association study workflow used to map key genes linked to important cattle traits.

Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.

Animals

DURABLE: A Workflow for Determining Corrosion-Driving and Protective Microbial Mechanisms.

Microbiologically influenced corrosion (MIC) threatens global infrastructure, causing billions of dollars in annual losses. Its persistence stems from unresolved mechanisms─particularly the metabolites produced by microorganisms that drive or inhibit corrosion─and the microbial community structures. Progress has been hindered by the absence of systematic workflows to rapidly and accurately identify MIC-relevant microorganisms and their functions. Here, we present DURABLE (Detection of Unique Corrosion Resistant or Accelerating Biologics in a Laboratory Environment), a pipeline that couples high-throughput microbial screening with genomic and metabolic workflows. We applied the DURABLE workflow to six diesel tank samples and revealed fuel-dependent microbial community structures, which showed greater diversity and evenness in bacterial communities than their fungal counterparts. The workflow used carbon steel beads to rapidly screen over 80 bacterial isolates for corrosive activity, reducing assay time to approximately 2 days compared with the conventional 30-day metal coupon test. More than 40 isolates were identified as corrosive. Further testing using mass spectrometry analysis revealed corrosion-associated metabolites, which were further validated using electrochemical assays. Thus, DURABLE achieved a ∼15-fold increase in screening speed and provided a scalable and mechanistic framework for dissecting MIC dynamics. We expect this advance will enable the development of precision mitigation strategies in hydrocarbon fuel infrastructure.

Bacteria

Benchmarking the OptiSpray-μPAC Workflow against a Traditional Nanospray Capillary Interface for Multiplexed Quantitative Proteomics.

Nanoflow liquid chromatography coupled with tandem mass spectrometry (LC-MS/MS) underpins modern quantitative proteomics, yet the column-to-mass spectrometer interface remains an important yet often underappreciated determinant of analytical depth, sensitivity, and reproducibility. Here, we benchmark an integrated workflow comprising the newly developed OptiSpray ion source and a micropillar array column (μPAC) cartridge against a conventional Nanospray Flex Source with an Accucore resin-packed capillary column. We performed a TMTpro 18-plex experiment across nine human cell lines on a FAIMS Pro-equipped Orbitrap Exploris 480. Following basic-pH reversed-phase fractionation, 12 fractions were analyzed on both workflow configurations under matched chromatographic gradient and acquisition conditions. Across both configurations, we quantified >9000 protein groups with highly comparable quantitative reproducibility and principal component clustering. Direct comparison of protein abundance ratios across cell lines showed agreement (Pearson R2 ≈ 0.7-0.8) without systematic bias. These results were achieved without workflow-specific optimization of the OptiSpray-μPAC platform, enabling direct transfer of established acquisition methods. Despite differences in column architecture, both configurations delivered comparable proteome coverage and quantitative fidelity. These findings establish the OptiSpray-μPAC workflow as a standardized alternative to conventional capillary-based interfaces, offering simplified operation while preserving quantitative performance.

Humans

Misdetection of frameshifts in SARS-CoV-2 genomes: need for additional harmonisation and efficient monitoring of data workflows.

Five years after the outbreak of the SARS-CoV-2 pandemic in 2020, diagnostic laboratories have moved from massive sequencing of thousands of samples to routine surveillance of SARS-CoV-2 cases, as with all other respiratory viruses. Surveillance remains of paramount importance to prevent a further SARS-CoV-2 surge, as the virus has been shown to mutate rapidly and can render available drugs and vaccines ineffective. During the pandemic, several bioinformatics pipelines and workflows have been developed to streamline analysis, shorten turnaround time and ensure reproducibility. As the number of samples decreases, laboratories are moving towards more flexible sequencing strategies and optimizing the cost per sample. However, workflow redesigns, even if individual steps have proven successful time and time again, can lead to challenges when changes in a bioinformatics pipeline are introduced (e.g. version updates, implementation of new features, etc.), a new combination of viral mutations emerge or a change in wet-lab procedures leads to unpredictable results. Here, we present a report of misidentified frameshift mutations in the consensus sequence of SARS-CoV-2, which led to an incorrect assumption of mutations in the spike and nucleocapsid viral proteins with the potential to affect PCR detection or even antigen testing. This investigation exemplifies the need for better awareness of the challenges that can occur even when using routinely applied protocols and analytical workflows and highlights the need for cooperation between experts of NGS, bioinformaticians and decision-makers towards more harmonized data workflows.

SARS-CoV-2

Parallel Analysis of Repeat Expansions: An Updated Clinical Nanopore Cas9-Targeted Sequencing Workflow for Nanopore R10 Flow Cells.

Hereditary ataxias, caused by expansions of short tandem repeats, are difficult to diagnose using traditional PCR and Southern blot methods, which struggle to detect complex repeat expansions and cannot assess repeat interruptions or methylation. An updated Clinical Nanopore Cas9-Targeted Sequencing workflow is presented for analyzing repeat expansions, now compatible with the Oxford Nanopore Technologies R10 flow cell. The workflow incorporates the Oxford Nanopore Technologies wf-human-variation Epi2Me workflow, including the Straglr tool to analyze base-called reads, ensuring compatibility with past, current, and future sequencing chemistries. It expands the number of genes analyzed from 10 to 27 and introduces new gene panels for ataxia, myopathy, neurodegeneration, and amyotrophic lateral sclerosis/motor neuron disease. Validated with Coriell reference and clinical samples, this method improves the analysis of pathogenic repeat expansions, providing deeper insights into repeat structures while addressing the limitations of traditional approaches. In this work, the use of multiplexing, Flongle flow cells, and single-gene targeting were explored as alternatives to panel-based approaches in the Clinical Nanopore Cas9-Targeted Sequencing workflow, finding that only single-gene targeting provides compatibility and reliable performance.

Journal Article

A modular class-aware workflow for small RNA sequencing analysis using mouse sperm as a case study.

BACKGROUND: Small RNA sequencing analysis is challenging because RNA classes differ in biogenesis, sequence redundancy, genomic organization, and annotation reliability. Integrated workflows accommodating these constraints remain limited, particularly for fragment-level and cluster-level analysis. METHODS: We present a reproducible, containerized, class-aware workflow for small RNA sequencing analysis, using mouse sperm as a case study. The workflow combines standardized preprocessing with complementary annotation and quantification strategies for microRNAs (miRNAs), transfer RNA-derived small RNAs (tsRNAs), ribosomal RNA-derived small RNAs (rsRNAs), and PIWI-interacting RNA (piRNA)-enriched genomic clusters. Using sperm small RNA data from offspring of lipopolysaccharide (LPS)-exposed male mice, we compared integrated-reference mapping, multi-class annotation, fragment-level tsRNA profiling, and genome-based piRNA cluster analysis, with custom modules for locus-aware harmonization and condition-specific cluster analysis. RESULTS: Integrated-reference mapping aligned 88.17% of reads and retained 690 features after filtering. It identified 11 differentially expressed miRNAs between LPS and controls, while other classes showed limited signal. Fragment-level profiling improved tsRNA resolution. piRNA cluster analysis identified 958 control and 940 LPS clusters, with 18 control-specific and no LPS-specific clusters. CONCLUSION: This workflow supports transparent, reproducible, class-aware interpretation of small RNA sequencing data while emphasizing cautious interpretation of piRNA-enriched signals from total small RNA sequencing.

Small non-coding RNA analysis

Triage and workflow optimization with artificial intelligence in pediatric imaging.

Artificial intelligence (AI) is being increasingly utilized in various aspects by the radiology department. With an ever-increasing burden on the healthcare system, particularly in emergency units, the need to incorporate AI in patient triage and workflow optimization cannot be overstated. Machine learning (ML)-based algorithms form the core of AI-based software, aiding healthcare professionals at nearly every step in delivering appropriate patient care. Regarding the radiology section of the hospital, AI-based algorithms have proven exceptionally useful in assisting radiologists and technicians with image acquisition. From accurate clinical referrals to scheduling computed tomography/magnetic resonance imaging scan appointments, from ensuring the lowest radiation exposure to offering timely follow-up reminders, ML-based software has indeed revolutionized the concept of modern image acquisition, especially in the pediatric radiology section. Although the implementation of these algorithms is swift, several technical challenges and the limited availability of pediatric datasets preclude their widespread use. The utility of multimodal pediatric datasets, which combine imaging, genomics, and clinical data, for comprehensive AI triage models can help AI systems evolve toward greater adaptability and integration, resulting in enhanced efficiency, reduced turnaround times, and improved patient outcomes in pediatric radiology departments in the future. In this article, we highlight and review the utility of AI and machine learning-based algorithms in efficiently aiding triage and streamlining the workflow in the pediatric radiology section, thereby ensuring an overall improvement in the departmental workflow.

Triage

optiPRM: A Targeted Immunopeptidomics LC-MS Workflow With Ultra-High Sensitivity for the Detection of Mutation-Derived Tumor Neoepitopes From Limited Input Material.

Personalized cancer immunotherapies such as therapeutic vaccines and adoptive transfer of T cell receptor-transgenic T cells rely on the presentation of tumor-specific peptides by human leukocyte antigen class I molecules to cytotoxic T cells. Such neoepitopes can for example arise from somatic mutations and their identification is crucial for the rational design of new therapeutic interventions. Liquid chromatography mass spectrometry (LC-MS)-based immunopeptidomics is the only method to directly prove actual peptide presentation and we have developed a parameter optimization workflow to tune targeted assays for maximum detection sensitivity on a per peptide basis, termed optiPRM. Optimization of collision energy using optiPRM allows for the improved detection of low abundant peptides that are very hard to detect using standard parameters. Applying this to immunopeptidomics, we detected a neoepitope in a patient-derived xenograft from as little as 2.5 × 106 cells input. Application of the workflow on small patient tumor samples allowed for the detection of five mutation-derived neoepitopes in three patients. One neoepitope was confirmed to be recognized by patient T cells. In conclusion, optiPRM, a targeted MS workflow reaching ultra-high sensitivity by per peptide parameter optimization, makes the identification of actionable neoepitopes possible from sample sizes usually available in the clinic.

Humans

Colora: a Snakemake workflow for complete chromosome-scale de novo genome assembly.

MOTIVATION: De novo assembly creates reference genomes that underpin many modern biodiversity and conservation studies. Large numbers of new genomes are being assembled by labs around the world. To avoid duplication of efforts and variable data quality, we desire a best-practice assembly process, implemented as an automated portable workflow. RESULTS: Here, we present Colora, a Snakemake workflow that produces chromosome-scale de novo primary or phased genome assemblies complete with organelles using Pacific Biosciences HiFi, Hi-C, and optionally Oxford Nanopore Technologies reads as input. Colora is a user-friendly, versatile, and reproducible pipeline that is ready to use by researchers looking for an automated way to obtain high-quality de novo genome assemblies. AVAILABILITY AND IMPLEMENTATION: The source code of Colora is available on GitHub (https://github.com/LiaOb21/colora) and has been deposited in Zenodo under DOI https://doi.org/10.5281/zenodo.13321576. Colora is also available at the Snakemake Workflow Catalog (https://snakemake.github.io/snakemake-workflow-catalog/? usage=LiaOb21%2Fcolora).

Software

Managing workflow executions with WESkit.

SUMMARY: In biomedical research, managing computational workflows across numerous projects-with varying parameters, tools, and environments-creates major challenges in scalability, reproducibility, and collaboration. Here, we present WESkit, an implementation of the Global Alliance for Genomics and Health (GA4GH) Workflow Execution Service (WES) interface, designed to streamline the execution, monitoring, and documentation of data processing workflows. It addresses the complexities involved in managing numerous executions with varying parameters across diverse research projects. Supporting both Snakemake and Nextflow, the system enables consistent automation and centralized monitoring, which benefits research groups aiming for long-term reproducibility and scalable collaboration. Its suitability for larger teams and service units is further enhanced by seamless integration into cloud environments, contributing to the GA4GH cloud framework. AVAILABILITY AND IMPLEMENTATION: The software WESkit is available under MIT license at the GitLab repository (https://gitlab.com/one-touch-pipeline/weskit). The WESkit main repository is archived at Software Heritage (https://archive.softwareheritage.org/browse/origin/directory/?origin_url=https://gitlab.com/one-touch-pipeline/weskit/api.git) and can be found using "one-touch-pipeline/weskit" term in the search section.

Workflow

Tractor workflow: a scalable Nextflow framework for local ancestry-aware genome-wide association studies.

MOTIVATION: The routine exclusion of admixed individuals from traditional genome-wide association studies (GWAS) due to concerns about spurious associations has limited multi-ancestry genetic discovery. Tractor addresses this issue by incorporating local ancestry into association testing, enabling the identification of ancestry-enriched signals and generating ancestry-specific summary statistics. However, adoption has been constrained by the complexity of prerequisite steps, including phasing and local ancestry inference, which require substantial bioinformatics expertise and introduce key analytical decision points. RESULTS: We developed a scalable, automated Nextflow workflow that integrates phasing, local ancestry inference, and Tractor association testing into a reproducible end-to-end pipeline. To demonstrate its utility, we applied the workflow to 32 blood biomarkers in 6245 two-way African-European admixed individuals from the UK Biobank. This pipeline performed efficiently at scale, replicating known associations and uncovering key ancestry-specific loci. These associations were largely driven by variants present on African ancestral tracts but absent from European tracts, underscoring the value of local ancestry-aware methods in uncovering previously masked genetic signals. AVAILABILITY AND IMPLEMENTATION: The workflow is modular, customizable, and compatible with commonly used phasing and local ancestry tools, minimizing manual intervention while preserving analytical flexibility. By lowering technical barriers to implementation, this framework facilitates broader adoption of local ancestry-aware GWAS, paving the way for expanded genetic discovery.

Humans

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes.

BACKGROUND: Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. METHODS: scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. RESULTS: We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. CONCLUSIONS: scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Single-Cell Gene Expression Analysis

A Robust, Self-Digestion-Resistant LysN with Superior Activity and Cleavage Fidelity for Advanced Proteomic Workflows.

LysN is a valuable protease in proteomics because it cleaves peptide bonds N-terminal to lysine, generating peptides with physicochemical properties complementary to those produced by LysC and trypsin. However, the broader adoption of LysN in proteomic workflows has been limited by the lack of commercially available enzymes that combine high activity, low missed-cleavage rates, and sufficient stability under practical sample-processing conditions. Here, we report the recombinant production and proteomic characterization of a self-digestion-resistant and highly active LysN from Shewanella loihica (SL-LysN). Using terminomics, we mapped the mature N- and C-termini of the enzyme and established the primary structure of the active protease. We further developed a high-density fermentation, refolding, and purification workflow to obtain highly purified recombinant SL-LysN. Biochemical and proteomic benchmarking showed that SL-LysN displayed 3.3-fold higher specific activity than commercial LysN and reduced missed cleavages by approximately 80%. Notably, SL-LysN retained high activity in the presence of 8 M urea or 1% SDS and showed strong resistance to autolysis, indicating exceptional robustness for proteomic sample preparation. In complex mammalian proteome digests, SL-LysN achieved >95% cleavage specificity and a missed-cleavage rate of only 5.9%. These features address a long-standing bottleneck in N-terminal proteolysis and establish SL-LysN as a high-performance enzymatic tool for advanced proteomic workflows, including deep protein sequencing, quantitative proteomics, terminomics, de novo sequencing and analyses requiring efficient digestion under denaturing conditions.

Shewanella