PubMed HealthSearch

SEARCH · PubMed Health

Results for “Automation pipeline”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

ALPAR: automated learning pipeline for antimicrobial resistance.

SUMMARY: The field of machine learning in antimicrobial resistance (AMR) research has experienced rapid growth, fueled by advancements in high-throughput genome sequencing and the growing capacity of computational resources. However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially for newcomers to the domain. In response to these challenges, we introduce ALPAR (Automated Learning Pipeline for Antimicrobial Resistance), a comprehensive AMR data analysis tool covering the entire process from processing of raw genomic data to training machine learning models to interpretation of results. Our method relies on a reproducible pipeline that integrates widely used bioinformatics tools, presenting a simplified, automatic workflow specifically tailored for single-reference AMR analysis. Accepting genomic data in the form of FASTA files as input, ALPAR facilitates the generation of machine learning-ready data tables and both the training of machine learning and the execution of genome-wide association studies (GWAS) experiments. Additionally, our tool offers supplementary functionalities such as phylogeny-based analysis of the distribution of mutations, enhancing its utility for researchers. The tool has also proven its performance in competitive benchmarks, winning the 2024 CAMDA Anti-Microbial Resistance Prediction Challenge and placing third in the 2025 edition. AVAILABILITY AND IMPLEMENTATION: ALPAR is open-source and freely accessible via GitHub (https://github.com/kalininalab/ALPAR). The pipeline is fully reproducible and can be easily installed as a Conda package (https://anaconda.org/kalininalab/ALPAR).

Machine Learning

Whole-genome automated assembly pipeline for Chlamydia trachomatis strains from reference, in vitro and clinical samples using the integrated CtGAP pipeline.

Whole genome sequencing (WGS) is pivotal for the molecular characterization of Chlamydia trachomatis (Ct)-the leading bacterial cause of sexually transmitted infections and infectious blindness worldwide. Ct WGS can inform epidemiologic, public health and outbreak investigations of these human-restricted pathogens. However, challenges persist in generating high-quality genomes for downstream analyses given its obligate intracellular nature and difficulty with in vitro propagation. No single tool exists for the entirety of Ct genome assembly, necessitating the adaptation of multiple programs with varying success. Compounding this issue is the absence of reliable Ct reference strain genomes. We, therefore, developed CtGAP-Chlamydia trachomatisGenome Assembly Pipeline-as an integrated 'one-stop-shop' pipeline for assembly and characterization of Ct genome sequencing data from various sources including isolates, in vitro samples, clinical swabs and urine. CtGAP, written in Snakemake, enables read quality statistics output, adapter and quality trimming, host read removal, de novo and reference-guided assembly, contig scaffolding, selective ompA, multi-locus-sequence and plasmid typing, phylogenetic tree construction, and recombinant genome identification. Twenty Ct reference genomes were also generated. Successfully validated on a diverse collection of 363 samples containing Ct, CtGAP represents a novel pipeline requiring minimal bioinformatics expertise with easy adaptation for use with other bacterial species.

Chlamydia trachomatis

SURE-Pipe: a pipeline to compare genomes and extract shared and unique regions.

Identification of unique and shared genomic regions between organisms has substantial translational potential for the development of marker-based diagnostic assays and sequence homology-driven taxonomic classification. An automated pipeline capable of performing genome comparisons at both the intra- and inter-species levels with minimal computational requirements can significantly advance genome-driven translational research. Species-specific genomic regions are particularly valuable for sequence-based species identification and for developing DNA amplification- or hybridization-based diagnostic assays. Here, we present SURE-Pipe, an automated and flexible pipeline for genome comparison and extraction of unique and shared genomic regions (https://github.com/BPaul-bioinfoLAB/SURE-Pipe). Benchmarking of this pipeline using simulated datasets demonstrated high accuracy for shared and unique region identification. Using the pairwise genome comparison module, six genome pairs from diverse microorganisms were analysed, and identified the unique and shared regions. In addition, the multigenome comparison module was applied to 96 genomes representing 24 Bacillus species and identified species-specific genomic regions. These regions were highly conserved among four strains of a species (>98% sequence identity) and exhibit little to no similarity with other species. Species-specific primers designed for all 24 Bacillus species showed no off-target amplification in in-silico polymerase chain reaction analysis, indicating their specificity. Overall, SURE-Pipe provides a robust and multipurpose framework for comparative genomics, and the outcomes can be used for species identification and the development of genome-based diagnostic approaches.

Genome, Bacterial

PULPO: pipeline of understanding large-scale patterns of oncogenomic signatures.

SUMMARY: PULPO v1.0 is a novel; fully automated pipeline designed for the preprocess and extraction of mutational signatures from raw Optical Genome Mapping (OGM) data. Built using Snakemake and executed within an isolated, Conda-managed environment, PULPO transforms complex cytogenetic alterations, captured at ultra-high resolution, into Catalogue of somatic mutations in cancer mutational signatures (COSMIC). This innovative approach not only enables researchers to work directly from raw OGM inputs but also streamlines the traditionally complex process of signature extraction, making advanced oncogenomic analyses accessible to users with varying levels of bioinformatics expertise. By facilitating the integration of comprehensive structural variants (SVs) and copy number variants (CNVs) data with established signature catalogues, PULPO paves the way for improved diagnostic accuracy and personalized therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The pipeline is open source and freely available under the MIT License at https://github.com/OncologyHNJ/PULPO-v.1.0 and DOI in Zenodo: https://zenodo.org/records/17749097.

Software

CoSAG-nf: A Scalable Nextflow Pipeline for Co-assembly, Optimization, and Interactive Visualization of High-Throughput Single-Cell Genomes.

MOTIVATION: Single-cell amplified genomes (SAGs) are crucial for resolving intra-population microbial heterogeneity and accurately understanding the metabolic potential of microbial dark matter populations. However, SAGs generated through multiple displacement amplification (MDA) of genomic DNA from single cells with single-copy chromosomes are highly fragmented and prone to contamination, severely hindering high-quality genome reconstruction and functional analysis, which greatly limits their scientific utility. Co-assembly of related SAGs can substantially improve genome quality, but to our knowledge no automated pipeline exists for high-throughput processing, forcing manual implementation of complex workflows that scale poorly to modern dataset sizes. RESULTS: We present CoSAG-nf, an automated high-throughput co-assembly and optimization pipeline for SAGs, implemented following the nf-core framework standards. The pipeline performs alignment-free clustering using sourmash MinHash signatures, then employs iterative tetranucleotide frequency profiling to identify and exclude outlier SAGs from co-assembly groups. CheckM2 quality assessment guides dynamic selection of optimal SAG combinations to optimize genome completeness and minimize contamination. Fully containerized, CoSAG-nf ensures reproducibility and scalability for the high-throughput processing of large-scale SAG datasets across diverse computing environments, including HPC and cloud platforms. The pipeline generates comprehensive HTML reports with quality metrics and taxonomic annotations, providing an end-to-end solution for automated high-throughput single-cell genome reconstruction. AVAILABILITY: CoSAG-nf is freely available under the MIT License at: https://github.com/linfengxu/CoSAG-nf. Archival code repository snapshots are published at zenodo with doi: https://doi.org/10.5281/zenodo.21525244. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Journal Article

Automated segmentation and length measurement of metacarpal and phalangeal bones for hand radiograph evaluation.

Evaluating hand and wrist radiographs is essential in pediatric endocrinology and clinical genetics, particularly for the assessment of suspected skeletal anomalies. In this study, we present Auto-Bone-Caliper, an automated system for the segmentation and length measurement of metacarpal and phalangeal (M&P) bones, trained and evaluated on public datasets comprising both normal and dysmorphic cases. We first introduce InstanceSAM, a two-stage framework that detects and segments all 19 M&P bones in pediatric hand radiographs, achieving Dice scores of 98.7% for normal bones and 95.0% for dysmorphic bones. We further develop and evaluate three methods for bone-length estimation, identifying a k-means-based approach as the most accurate, with relative errors of 2.2% for normal bones and 4.5% for dysmorphic bones. Our automated pipeline, Auto-Bone-Caliper, integrates InstanceSAM with the k-means-based length-estimation method. To enable scale-independent downstream analyses, we derive relative bone-length measures from the automated measurements. Using these relative measures, we statistically compare measurements obtained using Auto-Bone-Caliper on an independent dataset with a healthy reference catalog of normal bone morphologies, observing a high level of agreement (Wasserstein-1 distance = 0.012). Finally, we demonstrate a potential clinical use case of Auto-Bone-Caliper by obtaining relative metacarpophalangeal pattern profiles for three genetic conditions, namely Turner syndrome, achondroplasia, and pseudohypoparathyroidism. Our results highlight the potential of the Auto-Bone-Caliper to streamline and standardize M&P length measurement, providing an objective and reproducible tool suitable for clinical application.

Humans

Bioinformatics pipeline for the systematic mining genomic and proteomic variation linked to rare diseases: The example of monogenic diabetes.

Monogenic diabetes is characterized as a group of diseases caused by rare variants in single genes. Like for other rare diseases, multiple genes have been linked to monogenic diabetes with different measures of pathogenicity, but the information on the genes and variants is not unified among different resources, making it challenging to process them informatically. We have developed an automated pipeline for collecting and harmonizing data on genetic variants linked to monogenic diabetes. Furthermore, we have translated variant genetic sequences into protein sequences accounting for all protein isoforms and their variants. This allows researchers to consolidate information on variant genes and proteins linked to monogenic diabetes and facilitates their study using proteomics or structural biology. Our open and flexible implementation using Jupyter notebooks enables tailoring and modifying the pipeline and its application to other rare diseases.

Humans

Metabolism and gene expression models for the microbiome reveal how diet and metabolic dysbiosis impact disease.

The gut microbiome plays a critical role in human health, spurring extensive research using multi-omic technologies. Although these tools offer valuable insights, they often fall short in capturing the complexity of microbial interactions that associate with disease onset, progression, and treatment. Thus, integration of multi-omics datasets with metabolic models is needed to predict associations between microbial activity and disease. Here, we automated the reconstruction of 495 metabolic and gene expression models (ME-models), overcoming the main limitation preventing the wide use of this approach. We integrated them with multi-omics data from patients with inflammatory bowel disease (IBD), identifying taxa associated with variations in amino acids, short-chain fatty acids, and pH in the gut of IBD patients. In general, this approach provides testable hypotheses of the metabolic activity of the gut microbiota, and the automated pipeline opens the opportunity to study microbial interactions in other biologically relevant settings using ME-models.

Humans

Improved cohesin HiChIP protocol and bioinformatic analysis for robust detection of chromatin loops and stripes.

Chromosome Conformation Capture (3 C) methods, including Hi-C (a high-throughput variation of 3 C), detect pairwise interactions between DNA regions, enabling the reconstruction of chromatin architecture in the nucleus. HiChIP is a modification of the Hi-C experiment that includes a chromatin immunoprecipitation (ChIP) step, allowing genome-wide identification of chromatin contacts mediated by a protein of interest. In mammalian cells, cohesin protein complex is one of the major players in the establishment of chromatin loops. We present an improved cohesin HiChIP experimental protocol. Using comprehensive bioinformatic analysis, we show that a dual chromatin fixation method compared to the standard formaldehyde-only method, results in a substantially better signal-to-noise ratio, increased ChIP efficiency and improved detection of chromatin loops and architectural stripes. Additionally, we propose an automated pipeline called nf-HiChIP ( https://github.com/SFGLab/hichip-nf-pipeline ) for processing HiChIP samples starting from raw sequencing reads data and ending with a set of significant chromatin interactions (loops), which allows efficient and timely analysis of multiple samples in parallel, without requiring additional ChIP-seq experiments. Finally, using advanced approaches for biophysical modelling and stripe calling we generate accurate loop extrusion polymer models for a region of interest and provide a detailed picture of architectural stripes, respectively.

Chromatin

Autocycler: long-read consensus assembly for bacterial genomes.

MOTIVATION: Long-read sequencing enables complete bacterial genome assemblies, but individual assemblers are imperfect and often produce sequence-level and structural errors. Consensus assembly using Trycycler can improve accuracy, but its lack of automation limits scalability. There is a need for an automated method to generate high-quality consensus bacterial genome assemblies from long-read data. RESULTS: We present Autocycler, a command-line tool for generating accurate bacterial genome assemblies by combining multiple alternative long-read assemblies of the same genome. Without requiring user input, Autocycler builds a compacted De Bruijn graph from the input assemblies, clusters and filters contigs, trims overlaps, and resolves consensus sequences by selecting the most common variant at each locus. It also supports manual curation when desired, allowing users to refine assemblies in challenging or important cases. In our evaluation using Oxford Nanopore Technologies reads from five bacterial isolates, Autocycler outperformed individual assemblers, automated pipelines, and other consensus tools, producing assemblies with lower error rates and improved structural accuracy. AVAILABILITY AND IMPLEMENTATION: Autocycler is implemented in Rust, open-source, and freely available at github.com/rrwick/Autocycler. It runs on Linux and macOS and is extensively documented.

Genome, Bacterial

Combining Annotation Software to Identify Orthologous Genes (CASIO) Provides a New Dataset of Orthologous Genes for Swallowtail Butterflies.

With the massive increase in genomic resources, it is becoming increasingly popular to analyse thousands of loci across many species. However, many of the available genomes are not annotated, which hinders an efficient search for orthologous protein-coding genes. Here, we aim to develop a semi-automated pipeline and compare four genomic annotation methods (BRAKER2, BUSCO, Miniprot and Scipio). Our results highlight the importance of integrating multiple annotation tools to optimise ortholog detection and improve genomic studies. Each annotation method showed different strengths. BRAKER2 annotated a substantial number of genes. BUSCO, despite limitations inherent to its reference database, identified a higher number of orthologs. Miniprot exhibited notable flexibility in accommodating diverse protein datasets, whereas Scipio successfully recovered a considerable set of genes that were not detected by the other tools. The combination of these tools allowed for more comprehensive ortholog detection. Taking advantage of this pipeline, we developed a comprehensive dataset of orthologous genes for swallowtail butterflies (Lepidoptera: Papilionidae), called Papilionidae_odb, which will facilitate future studies, especially for a non-model group with abundant genomic data and few transcriptomic resources. We tested Papilionidae_odb by inferring a robust phylogenetic framework for Leptocircini using 142 complete genomes, which improved branch support for some phylogenetic relationships, although challenges remained in resolving relationships within certain species groups, likely due to rapid radiations. Our results highlight the complementary nature of the annotation methods and suggest that combining these tools can yield more accurate results in genomic research. This approach was implemented in a Snakemake workflow called CASIO (Combining Annotation Software to Identify Orthologous genes) and can easily be applied to other non-model groups to improve genomic datasets in diverse taxa where transcriptomic resources are still limited.

Animals

LncRAnalyzer: a robust workflow for long non-coding RNA discovery using RNA-Seq.

Long non-coding RNA (lncRNA) is a major transcript category that lacks protein-coding capabilities, with relatively low abundance and complex expression patterns. Distinguishing lncRNAs from protein-coding genes is a complex process involving multiple filtering steps. We developed an automated pipeline named LncRAnalyzer featuring retrained models for 60 species. This workflow aims to reduce the likelihood of obtaining protein-coding or partial protein-coding transcripts during lncRNA identification by utilizing eight distinct approaches. We conducted a 10-fold cross-validation of the sorghum models and training sets with their standard ones and other approaches using real-life RNA-Seq datasets and known lncRNA and CDS sequences of sorghum. The results showed that the sorghum models and training sets were outperformed. The pipeline output comprises upset plots illustrating the number of lncRNA/NPCTs identified by the approaches, commonly identified lncRNA and their classes, NPCTs, and expression count tables. A feature-level comparison and benchmarking analysis of LncRAnalyzer with four existing pipelines, namely, LncPipe, LncEvo, lncRNA-Annotation, and Plant-LncPipe, demonstrated that LncRAnalyzer is more comprehensive, easier to implement, and accurate in lncRNA predictions. This workflow also ascertains lncRNA origins from various Transposable Elements (TEs) in plants using TE annotations from APTEdb [http://apte.cp.utfpr.edu.br/]. LncRAnalyzer is publicly available on GitLab [https://gitlab.com/nikhilshinde0909/LncRAnalyzer.git] for academic users.

RNA, Long Noncoding

Evaluating 12 automated, whole-genome sequencing analysis pipelines for Mycobacterium tuberculosis complex: a comparative study.

BACKGROUND: Reliance on complex, custom-built bioinformatics pipelines is a barrier to the implementation of whole-genome sequencing (WGS) of Mycobacterium tuberculosis in high-burden settings in some low-income and middle-income countries (LMICs). Automated analysis pipelines could address this inequity in access to WGS-based diagnostics and surveillance. This study aimed to systematically evaluate the performance and usability of publicly available WGS pipelines for M tuberculosis. METHODS: We identified automated M tuberculosis WGS analysis pipelines through searches of PubMed and GitHub from database inception up to Aug 31, 2024. Accuracy, cost, accessibility, and scalability were assessed for each pipeline. We evaluated the accuracy of genotypic drug susceptibility testing (gDST) using publicly available sequences with phenotypic susceptibility data for 12 antituberculosis drugs. We estimated pooled sensitivity and specificity for each pipeline, across all drugs, by conducting a bivariate meta-analysis, with random effects representing between-drug variability. Lineage classifications were compared, and a previously epidemiologically well-characterised dataset was used to compare measures of genomic relatedness. FINDINGS: Among 28 candidate pipelines, 16 were excluded as they were unmaintained and inexecutable. 12 pipelines (11 compatible with Illumina and four compatible with Nanopore), all free to use, were included for evaluation. Six pipelines processed and stored data remotely, but for five of these six, scalability was limited by the need to upload sequences through web portals. For local processing pipelines, scalability was dependent on substantial local computational resources, data storage capacity, and command-line interfaces that limited user-friendliness. Only one of six remote-processing pipelines removed human DNA sequences before server upload. gDST was similarly accurate across ten of 11 Illumina-compatible pipelines and three of four Nanopore-compatible pipelines. All pipelines classified the main lineages consistently, although there were differences at sublineage resolution. Outputs from three of four pipelines reporting genomic relatedness were compatible with commonly cited single nucleotide polymorphism difference thresholds. INTERPRETATION: Numerous automated analysis pipelines capable of enhancing equity in M tuberculosis WGS are available. Given the overall similarities between the pipelines evaluated in this study in terms of gDST performance, lineage classification, and genomic relatedness inference, non-functional attributes such as availability, accessibility, scalability, and privacy could represent the point of difference for prospective users in LMICs with a high burden of tuberculosis. FUNDING: The Rhodes Trust, Wellcome, Ellison Institute of Technology, and the UK National Institute for Health and Care Research Oxford Biomedical Research Centre.

Mycobacterium tuberculosis

Comparative essentialome analysis of six Pectobacteriaceae strains using the TNSEEK pipeline identifies conserved and strain-specific fitness determinants.

Transposon sequencing (Tn-seq) is a powerful technique for defining the essential genes required for bacterial survival. However, gene essentiality can vary significantly across taxonomic levels, and comparing large Tn-seq datasets from multiple strains presents considerable analytical challenges. To address this, we developed TNSEEK, a fully automated bioinformatics pipeline for the systematic and comparative analysis of Tn-seq experiments. We applied TNSEEK to analyse newly generated data for six soft rot Pectobacteriaceae strains, encompassing species from the Dickeya and Pectobacterium genera, grown in a rich medium. This approach identified a core essentialome of 225 genes, primarily involved in fundamental cellular maintenance, conserved across all 6 strains, a set comparable in size to that of the neighbouring Enterobacteriaceae family. Only a few genus-specific essential genes were found, highlighting interesting distinct metabolic capabilities between Dickeya and Pectobacterium genera. In striking contrast, we discovered a large variable essentialome comprising 181 strain-specific genes, many of which are of unknown function. A portion of these strain-specific essential genes are components of defence systems and prophage genomic regions. The unexpected essentiality of selected components of these modules is consistent with cellular dependency on cognate toxic, restriction or immunity functions encoded by defence-associated loci under the tested growth condition. Furthermore, a comparison with the Escherichia coli essentialome demonstrates that discrepancies in gene essentiality can often be attributed to differences in growth conditions, particularly temperature, as well as variations in genetic redundancy. In conclusion, the TNSEEK pipeline provides a reproducible framework for comparative analysis of mariner/Himar1 Tn-seq datasets across multiple strains.

Pectobacterium

Equity in genome sequencing for rare disease diagnosis: a cross-sectional analysis of data from the UK 100,000 Genomes Project.

BACKGROUND: Genome sequencing has improved rare disease diagnosis and is now part of routine clinical care in the National Health Service in England. Automated prioritisation pipelines narrow millions of variants per patient to a small subset for clinical review, a process that relies on allele frequency resources that do not fully represent human genetic diversity. We assessed ancestry-related differences in variant prioritisation and diagnostic outcomes in patients from the UK 100,000 Genomes Project. METHODS: We analysed 29,405 rare disease probands with genome sequencing and linked clinical outcomes data. We used multivariable regression to assess ancestry-related differences in the number of variants prioritised for clinical review, the proportion of prioritised variants that were recorded as diagnostic, and diagnostic yield. We also evaluated the use of ancestry-stratified allele frequency filters derived from an independent, diverse UK cohort (n = 33,724). FINDINGS: Compared with the European ancestry group, the East African group had nearly three times more variants prioritised for clinical review (IRR 2.77, 95% CI 2.33-3.29). Other non-European groups also had significantly higher counts. Diagnostic yield was similar across ancestry groups after adjustment (LRT p = 0.1650). Prioritised variants were less likely to be recorded as diagnostic in East African (OR 0.32, 95% CI 0.22-0.46), West African (0.47, 0.39-0.57), South Asian (0.65, 0.58-0.73), and Middle Eastern (0.68, 0.54-0.86) groups. Applying ancestry-stratified allele-frequency filters removed 3.1% of prioritised variants overall-24.3% in the East African group-without loss of diagnostic sensitivity, including 29.5% of recorded VUS in this group. INTERPRETATION: Differences in the likelihood of prioritised variants being recorded as diagnostic partly reflect limitations of current allele frequency resources, which use broad population groupings that mask within-group diversity. Increased representation of diverse ancestries in reference databases and better estimation of ancestry-appropriate allele frequencies will help reduce inefficiencies and improve equity in variant prioritisation for rare disease diagnosis. FUNDING: The UK Department of Health and Social Care and the EU's Horizon 2020 Research and Innovation Programme.

Humans

Assessment of the impact of manual curation in BioCyc.

INTRODUCTION: BioCyc is an extensive collection of databases of genomic and pathway information for microorganisms and model eukaryotes. These organismal databases integrate diverse biological data by combining computationally inferred information, data imported from other databases, and, for selected organisms, literature-based manual curation. This study investigates the magnitude and significance of annotation changes performed during the curation of 10 prokaryotic genomes to better understand the rate of erroneous annotations and the value of BioCyc curation. METHODS: We identified curation changes by finding cases where the annotation of the protein at the start of the curation process differed from its annotation at the end of the process. RESULTS: We found that across a sample of curated databases (n = 10), the annotation of 6,753, or 25.6% of the proteins in the pooled protein dataset (n = 26,126) were modified. Assessment of considerable sampling fractions of these proteins found that a median of 62% (mean of 52.9%) represented functionally informative name changes, rather than stylistic annotation changes. These results were then extrapolated to total proteins with name changes with uncertainty quantified via finite population correction, indicating that most Tier 2 Biocyc PGDBs received hundreds of functionally informative name changes during manual curation. On average 363, or13% (±5.4% SD) of the proteins encoded in each genome received functionally informative annotation changes, ranging from 5.3% (Streptococcus pneumoniae D39V) to 22.7% (Staphylococcus aureus NCTC 8325). DISCUSSION: These findings demonstrate a substantial improvement in the accuracy of manually curated BioCyc databases compared with automated annotation pipelines. This result is particularly impactful as the rate of downstream propagation of erroneous annotations across biological databases can significantly compromise scientific discovery.

annotation errors

A computational ontology framework for the synthesis of multi-level pathology reports from brain MRI scans.

BackgroundConvolutional neural network (CNN) based volumetry of MRI data can help differentiate Alzheimer's disease (AD) and the behavioral variant of frontotemporal dementia (bvFTD) as causes of cognitive decline and dementia. However, existing CNN-based MRI volumetry tools lack a structured hierarchical representation of brain anatomy, which would allow for aggregating regional pathological information and automated computational inference.ObjectiveDevelop a computational ontology pipeline for quantifying hierarchical pathological abnormalities and visualize summary charts for brain atrophy findings, aiding differential diagnosis.MethodsUsing FastSurfer, we segmented brain regions and measured volume and cortical thickness from MRI scans pooled across multiple cohorts (N = 3433; ADNI, AIBL, DELCODE, DESCRIBE, EDSD, and NIFD), including healthy controls, prodromal and clinical AD cases, and bvFTD cases. Employing the Web Ontology Language (OWL), we built a semantic model encoding hierarchical anatomical information. Additionally, we created summary visualizations based on sunburst plots for visual inspection of the information stored in the ontology.ResultsOur computational framework dynamically estimated and aggregated regional pathological deviations across different levels of neuroanatomy abstraction. The disease similarity index derived from the volumetric and cortical thickness deviations achieved an AUC of 0.88 for separating AD and bvFTD, which was also reflected by distinct atrophy profile visualizations.ConclusionsThe proposed automated pipeline facilitates visual comparison of atrophy profiles across various disease types and stages. It provides a generalizable computational framework for summarizing pathologic findings, potentially enhancing the physicians' ability to evaluate brain pathologies robustly and interpretably.

Humans

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software