PubMed HealthSearch

SEARCH · PubMed Health

Results for “High-throughput sequencing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Colora: a Snakemake workflow for complete chromosome-scale de novo genome assembly.

MOTIVATION: De novo assembly creates reference genomes that underpin many modern biodiversity and conservation studies. Large numbers of new genomes are being assembled by labs around the world. To avoid duplication of efforts and variable data quality, we desire a best-practice assembly process, implemented as an automated portable workflow. RESULTS: Here, we present Colora, a Snakemake workflow that produces chromosome-scale de novo primary or phased genome assemblies complete with organelles using Pacific Biosciences HiFi, Hi-C, and optionally Oxford Nanopore Technologies reads as input. Colora is a user-friendly, versatile, and reproducible pipeline that is ready to use by researchers looking for an automated way to obtain high-quality de novo genome assemblies. AVAILABILITY AND IMPLEMENTATION: The source code of Colora is available on GitHub (https://github.com/LiaOb21/colora) and has been deposited in Zenodo under DOI https://doi.org/10.5281/zenodo.13321576. Colora is also available at the Snakemake Workflow Catalog (https://snakemake.github.io/snakemake-workflow-catalog/? usage=LiaOb21%2Fcolora).

Software

Composition-on-composition regression analysis for multi-omics integration of metagenomic data.

MOTIVATION: Compositional data are frequently encountered in many disciplines, such as in next-generation sequencing experiments widely used in biomedical studies. Regression analysis with compositional data as either responses or predictors has been well studied. However, when both responses and predictors are compositional, the inventory of analysis tools is surprisingly limited, especially in the high-dimensional setting. Among the few existing methods, most of them rely on a log-ratio transformation to move compositional data from the simplex to real numbers. Yet, a serious weakness of these methods is their failure to handle the substantial fraction of zeroes observed in data collected from next-generation sequencing experiments. RESULTS: To investigate associations between two high-dimensional multi-omics compositions, we propose a composition-on-composition (COC) regression analysis method which does not require log-ratio transformations and hence can handle zeroes in the data. To account for high dimensionality, we estimate regression coefficients using a penalized estimation equation approach. Finally, inference procedures for COC regression are also proposed. Superior performance of COC is demonstrated through both comprehensive numerical simulations and case studies. AVAILABILITY AND IMPLEMENTATION: Source R codes to implement COC method is available at https://github.com/nrios4/COC.

Regression Analysis

OctopuSV and TentacleSV: a one-stop toolkit for multi-sample, cross-platform structural variant comparison and analysis.

MOTIVATION: Structural variants (SVs) influence gene regulation, disease progression, and diagnostics, yet integrating SV calls across platforms remains difficult due to inconsistent annotations, limited merging flexibility, and fragmented workflows. Ambiguous breakend (BND) annotations, which comprise many variant calls, are often discarded or misclassified, hindering variant characterization. Existing tools lack advanced merging operations essential for precise identification of disease-specific or somatic variants across samples or patient groups. Additionally, current SV analysis pipelines require extensive manual intervention and complex parameter tuning, compromising reproducibility and scalability. Addressing these gaps is crucial for improving the accuracy, interpretability, and clinical utility of SV analyses. RESULTS: We developed OctopuSV and TentacleSV to address these long-standing challenges in SV analysis. OctopuSV features a specialized BND correction module that converts ambiguous BND annotations into canonical SV types, recovering important variants that are often overlooked by existing tools. Additionally, it provides advanced set operations (difference, complement, custom-defined) that enable sophisticated variant filtering without programming expertise, critical for identifying tumor-specific SVs or variants unique to specific sample groups. TentacleSV completes our solution by automating the entire SV analysis process from raw sequencing data to high-confidence callsets, ensuring consistency and reproducibility across projects. Benchmarking across short-read and long-read platforms showed superior F1 score, complete SV type consistency compared to existing tools. Our framework enables experimental biologists and clinical researchers to perform sophisticated analyses ranging from cancer subtype-specific SV identification to multi-sample comparative studies without requiring specialized programming skills. AVAILABILITY AND IMPLEMENTATION: All codes are available at https://github.com/ylab-hi/OctopuSV; https://github.com/ylab-hi/TentacleSV.

Software

Columba: fast approximate pattern matching with optimized search schemes.

MOTIVATION: Aligning sequencing reads to reference genomes is a fundamental task in bioinformatics. Aligners can be classified as lossy or lossless: lossy aligners prioritize speed by reporting only one or a few high-scoring alignments, whereas lossless aligners output all optimal alignments, ensuring completeness and sensitivity. RESULTS: This paper introduces Columba, a high-performance lossless aligner tailored for Illumina sequencing data. Columba processes single or paired-end reads in FASTQ format and outputs alignments in SAM format. By utilizing advanced search schemes and bit-parallel alignment techniques, Columba achieves exceptional speed. Columba is available in two variants. The first, based on the bidirectional FM-index, prioritizes speed. The second, Columba RLC, uses run-length compression using a bidirectional move structure, significantly reducing memory usage for large, repetitive datasets like pan-genomes. Benchmarks on the human genome, as well as bacterial and human pan-genome datasets, demonstrate that Columba is much faster than existing lossless aligners and even competitive with lossy tools. We integrated Columba into the OptiType HLA genotyping pipeline, where it substantially reduced computational time while maintaining accuracy. These results position Columba as a versatile, state-of-the-art tool for high-sensitivity genomic analyses. AVAILABILITY AND IMPLEMENTATION: The source code of Columba is available at https://github.com/biointec/columba under AGPL license. Scripts to reproduce the benchmarks and analyses are available at https://doi.org/10.5281/zenodo.15849246.

Software

Identification of autosomal and sex chromosome aneuploidies using next generation sequencing.

MOTIVATION: Chromosomal abnormalities, referred to as aneuploidies, occur in approximately 0.3% of live births. While the majority of aneuploidies in humans are incompatible with life, well-characterized exceptions include Down syndrome (47,+21), Patau syndrome (47,+13), Edwards syndrome (47,+18), Turner syndrome (45,X0), Klinefelter syndrome (47,XXY), and triple X syndrome (47,XXX). These chromosomal alterations disrupt gene expression and cellular function, leading to genetic and developmental disorders. With the increasing adoption of next generation sequencing (NGS) in clinical diagnostics, this study aims to explore the potential use of NGS for aneuploidies detection. RESULTS: Using data derived from clinical exomes (CES) and whole exomes (WES) sequencing we have been able to detect autosomal as well as sex chromosome aneuploidies with high specificity. Moreover, we have also been able to identify mosaic aneuploidies proving the high sensibility of this methodological approach. Thus, we present NGS as a cost-effective first line approach to detect chromosomal aneuploidies in routine diagnostic practice. AVAILABILITY AND IMPLEMENTATION: Scripts are available at https://github.com/B-R-I-D-G-E/AneuploidiesStudies.

Humans

dcHiChIP: a comprehensive Nextflow-based pipeline for multiscale analysis of chromatin architecture from HiChIP data.

MOTIVATION: Despite the growing use of HiChIP to investigate protein-directed chromatin architecture, a comprehensive and reproducible pipeline for analysing these datasets-from raw reads to multiscale 3D genome features-remains lacking. Existing tools often focus on isolated components, such as loop calling or matrix generation, but fall short in integrating structural annotation, functional enrichment, and spatial modeling within a unified framework. To address this gap, we developed dcHiChIP, a modular, scalable Nextflow-based workflow that streamlines the analysis of HiChIP data, enabling both routine processing and in-depth exploration of chromatin organization and regulatory interactions. RESULTS: dcHiChIP enables robust and reproducible analysis of HiChIP datasets across multiple scales of chromatin architecture. It accepts raw sequencing data as input and generates high-quality loop calls, domain annotations, and 3D genome models. It also performs functional annotation and motif enrichment analyses. Applied to benchmark CTCF HiChIP datasets, dcHiChIP identifies major chromatin architectural features such as TADs/CCDs, A/B compartments, and chromatin stripes, and offers efficient, end-to-end execution with support for batch processing and workflow resumability. AVAILABILITY: dcHiChIP is publicly available on GitHub at https://github.com/SFGLab/dcHiChIP, with documentation at https://sfglab.github.io/dcHiChIP/. The software version used in this study is archived at Zenodo: https://doi.org/10.5281/zenodo.22030542.

Chromatin

Assessing the readiness of Oxford Nanopore sequencing for clinical genomics applications.

Long-read sequencing (LRS) technologies, namely, Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio), have emerged as promising solutions to overcome the limitations of short-read sequencing (SRS). Nevertheless, the still higher sequencing error rates compared with SRS, need for customized pipelines, rapidly updating software, and incipient scalability continue to present challenges for adopting ONT in standard clinical practice. Here we assess the performance of ONT (R9 and R10 chemistries) in comparison to Illumina and MGI across 17 well-characterized reference samples with 11 clinical variants representing nine different genetic diseases. To enable this, we have implemented a production-ready pipeline including SNV, indel, STR, SV, and CNV detection, alongside reporting key summary metrics to ensure high-quality data at the production sequencing level. Our results show high accuracy of ONT across SNVs (F-score 0.978-0.983) and SVs (F-score = 0.75) but still weaknesses across indels (F-score 0.659-0.758). However, we highlight that ONT accurately detected all four pathogenic indels as well as the performance improvement in exons and with the newer R10 chemistry. We further demonstrated the importance of long reads to detect clinically impactful variants such as a FMR1 pathogenic expansion, often misclassified by SRS as being in the premutation range. Our multiplatform analysis and Sanger validation uncovered a 1 bp error in the Coriell annotation for a cystic fibrosis-causing indel in GM07829. This work underscores the growing readiness of ONT for clinical applications, highlighting both its advancements and its potential for broader adoption in clinical genomics and large-scale operations.

Humans

Comparative evaluation of probe-capture and conventional metagenomic sequencing across multiple clinical sample types, with analysis of paired bronchoalveolar lavage fluid and blood samples.

Conventional metagenomic next-generation sequencing (mNGS) suffers from host nucleic acid interference and poor performance in low-biomass samples. Probe-capture metagenomic sequencing (PC-mNGS), which enriches microbial targets via hybridization probes, shows superior sensitivity but lacks systematic multi-sample evaluations. This study compared PC-mNGS and mNGS across diverse clinical specimens (bronchoalveolar lavage fluid [BALF], blood, cerebrospinal fluid [CSF]) and assessed the clinical utility of pathogen co-detection in paired BALF-blood samples from sepsis patients. A total of 282 samples (81 BALF, 141 blood, 25 CSF, 35 others) sequenced by both PC-mNGS and mNGS were analyzed. Additionally, 621 paired BALF-blood samples from sepsis patients with pulmonary infections were evaluated. PC-mNGS achieved higher pathogen detection rates (66.67% vs 57.10%, P = 0.000198) than mNGS, particularly in blood (66.67% vs 47.52%, P = 2.5 × 10⁻⁵). PC-mNGS detected more bacteria (19 species exclusive) and fungi (11 species exclusive) than mNGS. Viruses showed comparable detection. BALF and CSF exhibited high overall agreement (OPA: 96.30% and 88%, respectively), while blood had lower concordance (NPA: 54.05%, OPA: 70.92%). A total of 60.55% of BALF-positive samples (PC-mNGS) had co-detected pathogens in blood. Gram-negative bacteria (e.g., Klebsiella pneumoniae) and fungi (e.g., Candida albicans) showed higher blood co-detection rates than viruses. In this study, PC-mNGS detected more pathogens and showed a higher positivity rate than mNGS in blood samples. BALF sequencing data, particularly bacterial reads per million (RPM), may predict bloodstream co-detection, aiding in sepsis management. However, clinical validation and integration with traditional diagnostics are needed to confirm utility. This study highlights PC-mNGS as a promising tool for complex infections but underscores the need for rigorous multi-context validation.IMPORTANCEAccurate and rapid identification of pathogens is critical for effective treatment of severe infectious diseases, such as sepsis. This study demonstrates that probe-capture metagenomic sequencing (PC-mNGS) detected more pathogens in blood samples compared to conventional metagenomic sequencing, especially for bacterial and fungal infections. By analyzing paired lung and blood samples, we show that high pathogen levels in lung fluid may predict bloodstream infection, offering a potential early warning for clinicians. These findings support the use of PC-mNGS as a more sensitive diagnostic tool, which could lead to faster, more targeted therapies and better outcomes for patients with complex infections.

Humans

Role of ctDNA Tumor Fraction in Selecting Immunotherapy-Based Regimens in Advanced Non-Small Cell Lung Cancer.

PURPOSE: Immune checkpoint blockers (ICB) have transformed advanced non-small cell lung cancer (aNSCLC) treatment, but identifying patients who benefit from adding chemotherapy remains challenging, especially in PD-L1 &#x2265; 50%. PD-L1 is an imperfect biomarker, highlighting the need for better selection tools. EXPERIMENTAL DESIGN: Liquid biopsy (LBx) assessment was performed using hybrid capture-based next-generation sequencing of plasma cell-free DNA. LBx data, molecular profile, and clinicopathologic data were collected. The predictive and prognostic values of tumor fraction (TF) were assessed using a deidentified nationwide (US-based) NSCLC clinicogenomic database [Clinico-Genomic Database (CGDB)]. An independent cohort with aNSCLC from Gustave Roussy was used to validate the findings and to study the correlation of circulating tumor DNA (ctDNA) TF and total metabolic tumor volume and its molecular correlates. RESULTS: In the CGDB database (n = 965), elevated ctDNA TF was prognostic for worse outcomes on ICBs and, when &#x2265;5%, predictive of benefit from ICB + chemotherapy [HR for real-world progression-free survival 0.58 (0.41-0.82); P = 0.002]. The 5% cutoff for TF was validated in an independent cohort from Gustave Roussy. In 283 patients with paired PET scans, ctDNA TF correlated with metabolic tumor volume (rho = 0.46; P < 0.001) and was influenced by TP53/RB1 mutations. CONCLUSIONS: ctDNA TF integrates disease burden and biology. Patients with high ctDNA TF derive greater benefit from chemoimmunotherapy, supporting its use as a biomarker to guide treatment intensification.

Humans

A hybrid and cost-efficient barcoding strategy for full-length 16S rRNA gene nanopore sequencing of environmental samples.

BACKGROUND: Accurate species-level identification of bacteria in complex environmental samples is essential for applications in biotechnology, ecological monitoring, and clinical diagnostics. Short-read platforms such as Illumina frequently truncate the 16S rRNA gene, limiting taxonomic resolution. In this work, we applied Oxford Nanopore Technology (ONT) long-read sequencing to full-length 16S rRNA amplicon in samples from natural soil amended with lignocellulosic biomass and a simplified microbial community derived from cultures grown on selective and differential carboxymethyl cellulose (CMC)-based substrates, with the aim to evaluate the difference in performance between a real, complex community and a less complex system. To reduce consumable costs, we substituted the standard ONT Barcoding kits with an in-house hybrid barcoding workflow. Specifically, PacBio PCR-based barcoding protocol was used for sample indexing, followed by library preparation using the ONT Ligation Sequencing Kit. This simplified approach retained compatibility with MinION and Flongle flow cells and supported accurate downstream demultiplexing while lowering barcode costs substantially. Additionally, a new bioinformatic workflow tailored to ONT data was implemented. RESULTS: Overall, the hybrid protocol significantly reduced per-sample barcoding costs while preserving high sequencing quality and throughput. The sequencing run yielded over 5 Gb of quality-filtered data (Q-score &#x2265; 10). Furthermore, the new bioinformatic workflow allowed taxonomic assignment at the species level for 49.38% of annotated taxa, compared to just 4.59% using Illumina NovaSeq sequencing of the V3-V4 region. ONT also recovered 2.3 times more genera and 1.3 times more families. Although 16S rRNA gene sequencing often cannot distinguish between closely related species, particularly within taxonomically complex groups, in this work, full-length reads substantially improved both taxonomic resolution and database matching. CONCLUSIONS: These results show that full-length 16S rRNA sequencing with ONT, paired with a low-cost barcoding strategy, enhanced taxonomic resolution compared to short-read workflows. This approach also offers a scalable and cost-effective option for high-resolution microbiome profiling in research and applied settings.

RNA, Ribosomal, 16S

Integrated analysis of gut microbiota, serum metabolomics, and proteomics reveals novel associations with clinical symptoms in patients with cerebral infarction.

BACKGROUND: Cerebral infarction (CI) is a major cause of adult disability and mortality worldwide. Mounting evidence supports the critical role of the gut-brain axis in cerebrovascular disease progression. This study aimed to characterize the alterations in gut microbiota, serum metabolome, and serum proteome in patients with CI, and to identify multi-omics signatures associated with clinical symptoms. METHODS: A total of 20 CI patients and 20 healthy controls (HC) were enrolled. Fecal microbiota was profiled using 16&#xa0;S rRNA gene high-throughput sequencing. Serum metabolomics and proteomics were analyzed using ultra-high-performance liquid chromatography-tandem mass spectrometry (UPLC-MS/MS) and data-independent acquisition (DIA) proteomics, respectively. Spearman correlation and multi-omics integration were applied to explore the associations among microbiota, metabolites, proteins, and clinical indicators. RESULTS: CI patients displayed significant gut microbiota dysbiosis, with a markedly lower gut microbiota health index (GMHI) and higher microbiota disorder index (MDI) compared with HC (P&#x2009;<&#x2009;0.001). The genera g_norank_o_RF39 and Oxalobacter were significantly enriched in CI patients, whereas Clostridium_sensu_stricto_1 and Agathobacter were enriched in HC. Metabolomic analysis identified 445 differential metabolites, mainly involved in glycerophospholipid metabolism, phenylalanine metabolism, and caffeine metabolism. Proteomic analysis revealed 140 differentially expressed proteins linked to inflammatory responses, calcium signaling, and NF-&#x3ba;B signaling. Multi-omics integration showed that signature gut microbiota was strongly correlated (P&#x2009;<&#x2009;0.005) with key serum metabolites and proteins implicated in CI pathogenesis. CONCLUSIONS: This integrated multi-omics study revealed distinct gut microbiota, serum metabolomic, and proteomic alterations in CI patients. The microbiota-metabolite-protein regulatory axes provide novel insights into the gut-brain axis in CI and may serve as potential diagnostic biomarkers or therapeutic targets.

Humans

Exploring the use of machine and deep learning in genome-wide association studies: a comprehensive review.

The advent of high-throughput sequencing technologies has generated increasingly large and complex genomic datasets, necessitating analytical approaches capable of capturing high-dimensional and potentially nonlinear genetic interactions. This situation has significantly impacted the entire field of Genome-Wide Association Study (GWAS), whose primary goal is the identification of genomic traits and variants that are statistically associated with the risk of a disease. However, traditional GWAS methods may show reduced performance when applied to highly polygenic and nonlinear genetic architectures. Computational strategies from Artificial Intelligence (AI) and, in particular, from machine- and deep-learning may provide a powerful tool to overcome such limitations, especially by capturing nonlinear interactions and complex hidden regularities in large-scale data, which traditional GWAS approaches might overlook. To date, only a few approaches have been introduced and systematically assessed. In this review, we describe the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS. Particular attention will be devoted to key issues, such as the interpretability of methods and results, and the curse of dimensionality. More specifically, the review presents 30 methods designed to leverage AI in GWAS, as well as presenting a comprehensive set of evaluation metrics for their performance, also providing references to the most frequently used databases, and biobanks. Overall, this work may serve as a starting point for both dry- and wet-lab researchers, aiming to extract deeper insights from genomic data by moving beyond traditional linear additive assumptions, and leveraging large-scale datasets through AI-driven approaches.

Artificial intelligence

[Applications and Challenges of Deep Learning in Human Genome Research].

In recent years, the advent of high-throughput omics technologies has fueled an explosive growth in human genomic data. Uncovering the latent functions within this vast data has become a significant challenge in functional genomics research. While traditional statistical methods have proved successful for analyzing smaller-scale datasets in the past, they exhibit clear limitations in analytical efficiency and integrating multi-dimensional data, struggling to meet the escalating demands of contemporary genomic analysis. The introduction of deep learning (DL) technologies offers a novel paradigm for this field. This review systematically examines the advances in applying deep learning to human genomics research. Studies demonstrate that when ample labeled data is available, discriminative DL computational methods-such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs)-achieve high accuracy and efficiency in genomic variant discovery tasks. Furthermore, generative DL methods, particularly Large Language Models (LLMs) leveraging self-supervised pre-training strategies, effectively integrate complex genomic information and exhibit superior performance in functional genomic sequence annotation and gene regulation studies. This review also explores the application of LLMs in multi-omics data integration and prediction. Looking ahead, the continued accumulation of long-read sequencing and high-dimensional data is expected to enable DL technologies to integrate increasingly complex and heterogeneous genomic information, playing an increasingly crucial role in human genomics research.

Deep Learning

Quantifying prevalence and risk factors of HIV multiple infection in Uganda from population-based deep-sequence data.

People living with HIV can acquire secondary infections through a process called superinfection, giving rise to simultaneous infection with genetically distinct variants (multiple infection). Multiple infection provides the necessary conditions for the generation of novel recombinant forms of HIV and may worsen clinical outcomes and increase the rate of transmission to HIV seronegative sexual partners. To date, studies of HIV multiple infection have relied on insensitive bulk-sequencing, labor intensive single genome amplification protocols, or deep-sequencing of short genome regions. Here, we identified multiple infections in whole-genome or near whole-genome HIV RNA deep-sequence data generated from plasma samples of 2,029 people living with viremic HIV who participated in the population-based Rakai Community Cohort Study (RCCS). We estimated individual- and population-level probabilities of being multiply infected and assessed epidemiological risk factors using the novel Bayesian deep-phylogenetic multiple infection model (deep&#xa0;-&#xa0;phyloMI) which accounts for bias due to partial sequencing success and false-negative and false-positive detection rates. We estimated that between 2010 and 2020, 4.09% (95% highest posterior density interval (HPD) 2.95%-5.45%) of RCCS participants with viremic HIV multiple infection at time of sampling. Participants living in high-HIV prevalence communities along Lake Victoria were 2.33-fold (95% HPD 1.3-3.7) more likely to harbor a multiple infection compared to individuals in lower prevalence neighboring communities. This work introduces a high-throughput surveillance framework for identifying people with multiple HIV infections and quantifying population-level prevalence and risk factors of multiple infection for clinical and epidemiological investigations.

Humans

CCNA2 orchestrates the PI3K/AKT signaling axis to propel prostate cancer metastasis.

BACKGROUND: Prostate cancer (PCa) remains one of the most common malignancies in men, posing a persistent global burden in terms of both public health and socioeconomic costs. Although early detection is essential for improving patient outcomes, existing clinical tools, including prostate-specific antigen (PSA) screening, digital rectal examination, and transrectal ultrasound-guided biopsy, are hampered by suboptimal specificity and positive predictive value, resulting in frequent overdiagnosis and overtreatment of indolent lesions while missing a subset of aggressive tumors at an early stage. In this context, the rapid advancement of high-throughput omics technologies, coupled with sophisticated machine learning (ML) algorithms, provides a powerful computational framework to dissect high-dimensional genomic data, uncover latent gene expression signatures, and identify candidate biomarkers with superior discriminative performance over conventional clinicopathological parameters. Therefore, in this study, we sought to screen for crucial ML-based biomarkers associated with PCa, with a particular focus on systematically assessing the diagnostic and prognostic value of CCNA2. Leveraging large-scale transcriptomic cohorts from public repositories, we employed an ensemble of ML approaches to prioritize candidate genes and subsequently evaluated the diagnostic performance of CCNA2 through receiver operating characteristic curve analysis, as well as its prognostic utility via Kaplan-Meier survival estimation and multivariate Cox proportional hazards modeling. Our findings are anticipated to elucidate the molecular landscape of PCa and offer a promising biomarker candidate for early detection and risk stratification. METHODS: This study integrated single-cell RNA sequencing, bulk transcriptomic data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) repositories, immunofluorescence, and multiple ML algorithms with in vitro functional assays to evaluate CCNA2 expression, clinical relevance, and biological behavior in PCa. RESULTS: CCNA2 was linked to metastasis and poor prognosis. High CCNA2 expression significantly correlated with adverse survival outcomes, and knockdown of CCNA2 suppressed proliferation, migration, and invasion in PCa cell lines. Mechanistically, CCNA2 modulated the PI3K/AKT signaling pathway. An ML-based diagnostic model incorporating CCNA2 demonstrated high predictive accuracy across multiple validation cohorts. CONCLUSIONS: CCNA2 serves as a promising prognostic biomarker and therapeutic target in prostate adenocarcinoma, driving tumor progression potentially via the PI3K/AKT axis.

CCNA2

Tumor Mutational Landscape and Its Correlation With Histopathological Characteristics in Breast Cancer.

BACKGROUND/AIM: In breast cancer, knowledge of the associations between clinicopathologic characteristics, genetic changes, and subtype-specific patterns is expanding. This study investigated how pathological and clinical variables affect the actionability of Next Generation Sequencing (NGS)-based tumor molecular data. MATERIALS AND METHODS: 227 breast cancer patients referred to Genekor's laboratory for tumor molecular profile analysis were included in the study. Pathology records were used to assess critical clinicopathological features, including HER2, ER, PR, Ki67, grade, metastatic site, and age. A 1021-gene NGS-based multigene panel was utilized to assess tumor biology alongside tumor mutational burden (TMB) and microsatellite instability (MSI). RESULTS: Comprehensive genomic profiling revealed that 95.6% of the patients harbored at least one oncogenic or likely oncogenic alteration, highlighting the high diagnostic yield of NGS-based testing. Distinct subtype-specific patterns were observed: HR+/HER2- tumors were enriched for PIK3CA and ESR1 gene alterations, whereas triple-negative breast cancer (TNBC) was dominated by TP53 alterations. Clinically actionable alterations were most common in HR+/HER2- tumors (~60% on-label), whereas TNBC more often harbored off-label or trial-associated targets. The inclusion of tumor-agnostic biomarkers (TMB/MSI) increased on-label actionability up to 64.5% in HR+/HER2- tumors, primarily driven by TMB-high cases. Median TMB values were low, and age was the only independent predictor. Furthermore, the presence of actionable alterations was significantly higher in metastatic tumors, and TP53 alterations were associated with aggressive tumor characteristics. CONCLUSION: Comprehensive NGS-based genomic profiling identifies clinically actionable alterations in over half of breast cancer patients, with substantial variability across molecular subtypes. The HR+/HER2- subtype demonstrates the highest prevalence of on-label actionable biomarkers. These findings support the routine implementation of comprehensive genomic profiling, especially in metastatic HER2-negative breast cancer, to guide precision oncology strategies and enable enrollment in biomarker-driven clinical trials.

Humans

Lawsonella clevelandensis: a normal flora that bites deep.

Since its formal description in 2016, Lawsonella clevelandensis-a strictly anaerobic, partially acid-fast bacterium-has been increasingly recognized as a cause of deep-seated abscesses, yet its fastidious nature and absence from routine diagnostic databases contribute to significant underdiagnosis. This narrative review synthesizes current knowledge on its microbiology, expanding clinical spectrum, diagnostic strategies, and treatment, based on a literature search of PubMed and Web of Science up to April 2026. Analysis of 27 documented publications, comprising 18 clinical cases, reveals a potential association with host risk factors including diabetes, immunosuppression, and prior surgical procedures, alongside a notable predilection for fat-rich tissues such as the breast and abdomen. While metagenomic next-generation sequencing and 16S rRNA gene amplification have become indispensable for definitive identification, antimicrobial susceptibility data-derived primarily from a single strain-demonstrate uniformly low minimum inhibitory concentrations for penicillins, carbapenems, clindamycin, and metronidazole, with no acquired resistance genes identified by whole-genome sequencing. However, the absence of established clinical breakpoints and limited tested isolates precludes definitive conclusions about universal susceptibility. Clinicians should maintain a high index of suspicion for L. clevelandensis in culture-negative deep abscesses, particularly those with acid-fast rods, as prompt diagnosis and empirical therapy with &#x3b2;-lactam/&#x3b2;-lactamase inhibitors or carbapenems appear reasonable based on current in vitro and clinical evidence, though further susceptibility surveillance is essential.

Humans

Isolation of folate-producing probiotic candidates and their effects on homocysteine metabolism and gut microbiota composition.

BACKGROUND: Folate deficiency is a global nutritional problem associated with multiple adverse health outcomes, including impaired one-carbon metabolism and elevated homocysteine levels (hyperhomocysteinemia). Gut microbiota-mediated folate biosynthesis has emerged as a promising strategy for improving the host's folate status. This study aimed to isolate folate-producing probiotic strains, clarify their folate synthesis mechanisms, and evaluate their regulatory effects on folate metabolism and gut microbiota. METHODS: High-throughput cultivation and screening were performed to isolate folate-producing candidate probiotics. Whole-genome sequencing analysis, pathway reconstruction, and metabolite profiling in fermented milk were performed to explore folate biosynthesis pathways and microbial cross-feeding interactions. A folate-deficient mouse model was established to evaluate the effects of a candidate probiotic cocktail on serum folate, homocysteine (Hcy) levels, and gut microbiota composition using quantitative PCR (qPCR) and 16S rRNA gene sequencing. RESULTS: High-throughput screening identified 8 high-folate-producing candidate probiotic strains, including Lactiplantibacillus plantarum and Heyndrickxia coagulans, from over 1,000 isolates. Genomic analysis revealed that most commonly used probiotics lacked para-aminobenzoic acid (pABA) biosynthesis genes but retained downstream modules, suggesting a reliance on cross-feeding with pABA-producing gut commensals such as Bacteroides. Metabolite profiling of fermented milk demonstrated that selected strains significantly increased bioactive 5-methyltetrahydrofolate (5-MeTHF) and tetrahydrofolate levels. In vivo, only a high-dose candidate probiotic cocktail significantly elevated serum folate (p&#x202f;<&#x202f;0.05) and reduced homocysteine levels (p&#x202f;<&#x202f;0.05) in deficient mice. Fecal qPCR confirmed dose-dependent transient persistence of the administered bacterial species. Consistent with the qPCR data, 16S rRNA gene sequences demonstrated significant enrichment of these administered species observed in the high-dose group. Furthermore, beta-diversity analysis found that high-dose candidate probiotic supplementation promoted a shift in the gut microbiota composition toward a normal profile, partially mitigating the dysbiosis induced by the folate-deficient diet. This effect was accompanied by a significant enrichment of potential short-chain fatty acid producers (e.g., Lachnospiraceae and Oscillospiraceae) and the depletion of potential opportunistic pathogens. CONCLUSION: This study screened high-folate-producing candidate probiotic strains and demonstrated their ability to synthesize the active form of 5-MeTHF. Moreover, folate-producing candidate probiotic cocktail treatment significantly improved folate status and Hcy metabolism and modulated the gut microbiota by enriching potential beneficial bacterial taxa. These findings suggested that folate-producing probiotics may serve as a promising microbiota-based strategy to improve folate availability and homocysteine metabolism.

B vitamin