PubMed HealthSearch

SEARCH · PubMed Health

Results for “High-throughput sequencing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Differentiating tuberculous pleurisy from pulmonary tuberculosis using mNGS: a multicenter cohort analysis.

BACKGROUND: Tuberculous pleurisy (TBP), a major extrapulmonary form of tuberculosis, is characterized by a paucibacillary state that makes diagnosis challenging. Metagenomic next-generation sequencing (mNGS) has emerged as a promising approach for MTB detection; however, its discriminatory value between TBP and pulmonary tuberculosis (PTB) among mNGS-confirmed cases, and its integration with clinical features for differential diagnosis, remain insufficiently defined. METHODS: This multicenter retrospective cohort included hospitalized patients with MTB-positive mNGS results from January 2020 to January 2025. As only mNGS-positive cases were included, overall mNGS diagnostic sensitivity cannot be estimated. Twelve TBP patients were matched 1:2 with twenty-four PTB patients by age and sex; patients with immunosuppressive conditions were excluded prior to matching. Clinical, laboratory, mNGS, and conventional TB test data were collected. Logistic regression and ROC analyses were performed. RESULTS: Conventional tests showed limited sensitivity in TBP despite universal mNGS positivity. MTB read counts were similar between groups (median 1976.5 vs. 990.0, P = 0.920). Pleural-derived specimens predominated in TBP (41.7% vs. 4.2%, P = 0.007). CRP demonstrated the highest individual discriminatory value (AUC = 0.658, P = 0.131), though no single predictor reached significance. A combined model (cough, fever, CRP, WBC) showed modest non-significant improvement (AUC = 0.722, overall P = 0.359; sensitivity 66.7%, specificity 83.3%). Given EPV ≈ 3, all findings are exploratory only. No significant prognostic predictors were identified in TBP; a non-significant trend toward lower lymphocyte counts was observed in patients with unfavorable outcomes (0.60 vs. 1.10 ×109/L, P = 0.115). CONCLUSIONS: Among mNGS-confirmed cases, MTB read counts were comparable between TBP and PTB. No single parameter reliably distinguished the two; a combined clinical model showed modest improvement but requires prospective validation in larger cohorts. Integrating mNGS with systematic clinical evaluation remains essential for accurate TB diagnosis.

Humans

Detecting and reconstructing breakage-fusion-bridge cycles from long-read sequencing using BFBArchitect.

MOTIVATION: Focal oncogene amplification is a key driver of tumor progression. Remarkably, the increased pathology depends on the context-whether the amplification is extrachromosomal (ecDNA) or intrachromosomal. EcDNA amplifications promote heterogeneity, therapy resistance, and poor prognosis. Focal intrachromosomal amplifications often arise through breakage-fusion-bridge (BFB) cycles, which produce highly rearranged but stable chromosomes. Distinguishing BFB from ecDNA remains challenging due to overlapping genomic signatures. To address this, we present BFBArchitect, a computational method leveraging long-read Oxford Nanopore data to identify BFB sequences consistent with both copy number and structural variations. RESULTS: We provide a novel combinatorial characterization of BFB, which naturally leads to an integer linear programming (ILP) optimization. The ILP optimization generates a BFB sequence that best explains experimentally observed copy numbers and foldback structural variants. We implement this idea in a tool called BFBArchitect, which achieves near-perfect accuracy in distinguishing BFB from non-BFB structures in extensive simulations as well as on 18 validated tumor samples. Moreover, it generates sequence-level BFB reconstructions that provide mechanistic insights into BFB formation, including repair mechanisms with template switching and other structural variants, and recapture of telomere for stabilization. AVAILABILITY AND IMPLEMENTATION: BFBArchitect is available at https://github.com/AmpliconSuite/BFBArchitect.

Sequence Analysis, DNA

Accurate somatic small variant discovery for multiple sequencing technologies with DeepSomatic.

Somatic variant detection is an integral part of cancer genomics analysis. While most methods have focused on short-read sequencing, long-read technologies offer potential advantages in repeat mapping and variant phasing. We present DeepSomatic, a deep-learning method for detecting somatic small nucleotide variations and insertions and deletions from both short-read and long-read data. The method has modes for whole-genome and whole-exome sequencing and can run on tumor-normal, tumor-only and formalin-fixed paraffin-embedded samples. To train DeepSomatic and help address the dearth of publicly available training and benchmarking data for somatic variant detection, we generated and make openly available the Cancer Standards Long-read Evaluation (CASTLE) dataset of six matched tumor-normal cell line pairs whole-genome sequenced with Illumina, PacBio HiFi and Oxford Nanopore Technologies, along with benchmark variant sets. Across samples, both cell line and patient-derived, and across short-read and long-read sequencing technologies, DeepSomatic consistently outperforms existing callers.

Humans

Targeted sequencing reveals a distinct genetic alteration landscape in oral multiple primary squamous cell carcinomas.

OBJECTIVE: Oral multiple primary cancers (MPCs) are associated with poor clinical outcomes, yet their genomic characteristics remain insufficiently understood. DESIGN: Fifty-four formalin-fixed paraffin-embedded (FFPE) tumor samples from 30 patients with oral MPCs were analyzed using high-depth targeted sequencing of a customized 14-gene panel derived from prior whole-exome sequencing data. Detected alterations were analyzed after removal of synonymous mutations. RESULTS: Non-silent genomic alterations were identified in 59.3% (32/54) of samples, involving 19 patients. A total of 70 variant loci across 13 genes were detected. AKAP13 was the most frequently mutated gene at both the sample (22.2%, 12/54), with recurrent mutations observed across multiple patients. In contrast, TP53 mutations occurred at a substantially lower frequency (11.1%, 6/54). Marked inter- and intra-patient mutational heterogeneity was observed. CONCLUSIONS: FFPE-based targeted sequencing enabled an initial characterization of genomic alterations in oral MPCs. Recurrent alterations in AKAP13, GLI2, JMJD1C, and DNAH8, together with the relatively low frequency of TP53 alterations, identify candidate genomic features for further investigation and provide a basis for future studies of the molecular basis of oral MPCs.

Humans

Upscaling Genotyping by Amplicon Sequencing With GBAS-GUI.

Genotyping by amplicon sequencing (GBAS) is a relatively low-cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large-scale genetic monitoring projects. However, most existing analytical pipelines are either marker-specific, insufficiently scalable, or lacking efficient data management systems for the long-term integration of genotypic information, limiting the full potential of GBAS. Here, we address this gap by introducing GBAS-GUI (https://github.com/sonnenbe-dot/GBAS-GUI), a pipeline capable of generating GBAS-based genotypic data for a wide variety of loci at scale. GBAS-GUI integrates a graphical user interface with multiple checkpoints to improve accessibility and robustness. It implements multiprocessing architecture and a relational database that links genotypic data with associated sample metadata to enhance scalability and data management. The pipeline further enables marker screening through automated calculation of polymorphism information content (PIC) and implements a strategy to recover homologous genotypic information from paralogous loci with non-overlapping amplicon length ranges. Using multiple empirical datasets, we demonstrate substantial improvements in processing speed, database management and handling artefacts related to co-amplification of unspecific regions and duplicates of the same genomic region. We further show that incorporating the full sequence information captured by an amplicon increases marker information content beyond what is achievable with length-based genotyping alone and expands the analytical versatility of GBAS. Overall, GBAS-GUI provides a robust, scalable and versatile framework that unlocks the potential of GBAS for large-scale population genetic and phylogeographic studies.

Genotyping Techniques

Sequencing approaches in hereditary cancer testing: strengths, limitations and future directions.

Over the past three decades, Hereditary Cancer Testing (HCT) has evolved from single gene assays into multigene panel testing (MGPT), which allows for the screening of all known hereditary cancer genes in a single assay. MGPT is currently the standard approach for clinical HCT. However, with decreasing sequencing costs and increased instrument throughput, the scalability of exome sequencing (ES) and genome sequencing (GS) for HCT indications is becoming more viable. These methods provide broader insights into the coding exons and/or the entire genome, respectively. ES/GS data can also be reanalyzed to identify variants in novel genes that were not characterized at the time of initial testing, or to support research efforts aimed at uncovering additional associations between germline variants and cancer predisposition. Additionally, the emerging use of long-read sequencing (LRS) is noteworthy, enabling improved variant detection compared to short-read sequencing, especially for complex/structural variants and variation in difficult-to-sequence or paralogous regions in genes such as PMS2. This has the potential to increase the accuracy of HCT, reduce the turnaround time, find previously unidentifiable cancer risk variants, and ultimately increase the diagnostic yield. This article provides a comprehensive summary of the sequencing approaches used in HCT, discussing their strengths and limitations. We also highlight the added value of complementing DNA-only testing with RNA and tumor sequencing. Furthermore, we explore LRS-based approaches and discuss opportunities for their implementation in routine genetic testing for hereditary cancer.

Humans

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans

DirectASRM: uncovering allele-specific post-transcriptional RNA modifications through direct RNA sequencing.

SUMMARY: We developed DirectASRM, a comprehensive database for the systematic identification, integration, and annotation of allele-specific RNA modifications (ASRMs) from direct RNA sequencing data. DirectASRM enables single-base, transcript-level detection of ASRMs across multiple RNA modification types, diverse organisms and condition-specific contexts. The database further evaluates the confidence of each ASRM-SNP pair association within isoform context by jointly considering statistical evidence of allelic modification imbalance and independent support from external next-generation sequencing (NGS) - based RNA modification resources. DirectASRM also provides extensive functional annotations for ASRMs and their associated variants, including intra-sample transcript-level allele-specific expression (ASE) and allele-specific splicing, as well as additional post-transcriptional regulatory features such as miRNA binding, circRNA, RNA-protein interactions, and disease relevance. Overall, DirectASRM serves as a comprehensive resource that supports systematic investigation of the potential functional impact of genetic variants in epitranscriptomic regulation. AVAILABILITY AND IMPLEMENTATION: DirectASRM database is freely accessible at http://modinfor.com/DirectASRM/. DirectASRM pipeline is available at GitHub (https://github.com/jiayin1101/DirectASRM_pipeline) and Zenodo (DOI: https://doi.org/10.5281/zenodo.19876077).

Alleles

Analysis of deep-resequencing data of 984 soybean accessions reveals structural variations underlying agronomic traits.

Genomic structural variants (SVs) are major sources of genetic variation and have profound impacts on phenotypic traits. However, their functional effects remain largely unexplored in soybean. Here, we resequence 940 soybean accessions. Together with 44 publicly available datasets, we identify 602,281 SVs. Using a graph-based genome, we detect an additional 58,760 presence/absence variations (PAVs) that broadly affect gene expression. Population genomic analyses reveal that SVs serve as a core driving force for soybean domestication and improvement. Integrating SVs with QTLs for oil and protein content, and performing GWAS on 27 traits, we identify key functional SVs. These include transposable element insertions altering seed coat color, multiple insertions within a cytochrome P450 gene modifying flower and hypocotyl color, and a GmMATE1 deletion enhancing seed size. Together, our study establishes a comprehensive SV map of soybean, offering a valuable resource for dissecting the genetic basis of complex traits to accelerate molecular breeding.

Glycine max

Optimizing GRIDSS for clinical use: A targeted NGS filtering strategy for germline structural variant detection.

Detecting intermediate-sized structural variants (SVs) remains challenging in diagnostics, as tools for single-nucleotide and copy-number variants, particularly read-depth-based methods, are often insufficient. GRIDSS addresses this gap by integrating paired-end mapping, split-read analysis, and assembly-based approaches. However, its use in targeted sequencing and diagnostic workflows remains complex. NGS panel data from 9726 patients with suspected hereditary cancer were analyzed using GRIDSS. A filtering strategy was developed to prioritize clinically relevant germline SVs. Multiple parameter settings were tested to optimize performance. The initial dataset of 1,307,592 variants was reduced to 89 candidates after applying the selected filtering strategy. Of these, 24 had been previously detected by routine callers and were not further analyzed. Among the remaining 65, 13 were considered likely true positives after visual inspection using IGV. Experimental validation was performed by Sanger/Nanopore long-read sequencing for these variants, all of which were confirmed. Eight were classified as (likely) pathogenic, including two frameshift duplications in MSH6, one splicing variant in BARD1, and five mobile element insertions in APC, BRCA2, and PALB2. Altogether, GRIDSS implementation increased diagnostic yield while maintaining feasibility for diagnostic workflows. Comprehensive workflow scheme for germline structural variant detection and results in our diagnostic setting.

Humans

Lynch syndrome-associated urothelial carcinoma: clinical and molecular findings from a single-institution cohort.

Lynch syndrome-associated urothelial carcinoma (LS-UC) is a rare and undercharacterized clinical entity. While FGFR3 alterations are well described in sporadic urothelial carcinoma, their prevalence and clinical implications in LS-UC remain unclear. We aimed to provide a comprehensive clinical and molecular characterization of LS-UC. We conducted a retrospective single-center study including patients with Lynch syndrome (LS) and histologically confirmed urothelial carcinoma (UC). Clinical, pathological, treatment, and follow-up data were collected. Targeted next-generation sequencing was performed on available tumor samples to assess genomic alterations, with particular attention to FGFR3 mutations. A total of 27 patients with LS-UC were identified, with a predominance of upper urinary tract involvement (70%). Most tumors were diagnosed at an early stage and initially managed with local treatment. During a median follow-up of 92 months, 48% of patients experienced recurrence, with a median time to recurrence of 37 months. Recurrences were predominantly local and were mainly managed with additional surgical or intravesical treatments. No deaths were attributable to UC at last follow-up. Molecular analysis was feasible in 9 cases. FGFR3 mutations were detected in 67% of evaluable samples, with the recurrent p.Arg248Cys hotspot identified in 55% of cases. Additional alterations involved TP53, SWI/SNF complex genes, and PIK3CA, which co-occurred with FGFR3 p.Arg248Cys. No gene fusions were identified. This study expands the limited molecular and clinical evidence on Lynch syndrome-associated urothelial carcinoma. Beyond confirming the recurrent role of FGFR3 (notably p.Arg248Cys), our comprehensive multigene profiling enriches the current genomic knowledge for this rare population. Multi-center collaborative efforts remain essential to aggregate larger datasets and ultimately guide personalized patient management.

Humans

Obtaining a Diagnostic Yield via Scan findings prior to the introduction of SEquencing retrospectivelY (ODYSSEY): a cohort study.

OBJECTIVE: To determine the retrospective yield of prenatal exome sequencing (PES) by establishing the proportion of children with a postnatal monogenic diagnosis that could have been diagnosed prenatally if PES had been available. METHODS: The study cohort comprised a sample of children in Northern Ireland, born between January 2010 and January 2018 (predating routine availability of PES), who received a monogenic diagnosis postnatally via next generation sequencing as part of either of two UK-wide studies (the 100 000 Genomes Project (2015-2018) or the Deciphering Developmental Disorders study (2011-2015)). Clinical data were collected retrospectively and correlated with the current UK National Health Service PES protocol, including the phenotypic eligibility criteria for PES and the associated fetal anomalies gene panel. Cases were considered retrospective diagnoses if the fetal phenotype would have been eligible for PES and the diagnostic gene was included on the test panel, meaning prenatal diagnosis in this current era could have been feasible. RESULTS: Of 101 children, 17.8% (95% CI, 10.3-25.3%) had both an eligible fetal structural anomaly (FSA) (i.e. high-risk FSA) and a diagnostic gene on the associated test panel, meaning that they could have been diagnosed prenatally in the current clinical landscape. The median length of the diagnostic odyssey for this subgroup of children was 3.7 years (1354 (range, 822-2450) days). Moreover, 58.4% (n = 59) of cases had no anomalies detected prenatally and 19.8% (n = 20) had a FSA that would not meet the eligibility criteria for PES (low-risk FSA). Although these cases would have been ineligible for PES under the current clinical pathway, 89.9% (n = 71/79) were affected by severe or profound syndromes. Postnatally, the most common functional anomalies were neurodevelopmental delay/intellectual disability and/or behavioral abnormality, which were observed in 80.2% (n = 81) of the included children. However, 80.2% (n = 65/81) of these affected children did not present with fetal anomalies eligible for PES. CONCLUSIONS: Almost one-fifth of children with a monogenic condition included in this study could have received a diagnosis via modern PES, avoiding a diagnostic odyssey lasting almost 4 years. However, despite having a monogenic condition, over half of the children did not present with any structural anomalies in utero. This demonstrates the degree to which fetal imaging is limited in its ability to reassure parents of the absence of a fetal genetic syndrome. © 2026 The Author(s). Ultrasound in Obstetrics & Gynecology published by John Wiley & Sons Ltd on behalf of International Society of Ultrasound in Obstetrics and Gynecology.

Humans

Comparative evaluation of probe-capture and conventional metagenomic sequencing across multiple clinical sample types, with analysis of paired bronchoalveolar lavage fluid and blood samples.

Conventional metagenomic next-generation sequencing (mNGS) suffers from host nucleic acid interference and poor performance in low-biomass samples. Probe-capture metagenomic sequencing (PC-mNGS), which enriches microbial targets via hybridization probes, shows superior sensitivity but lacks systematic multi-sample evaluations. This study compared PC-mNGS and mNGS across diverse clinical specimens (bronchoalveolar lavage fluid [BALF], blood, cerebrospinal fluid [CSF]) and assessed the clinical utility of pathogen co-detection in paired BALF-blood samples from sepsis patients. A total of 282 samples (81 BALF, 141 blood, 25 CSF, 35 others) sequenced by both PC-mNGS and mNGS were analyzed. Additionally, 621 paired BALF-blood samples from sepsis patients with pulmonary infections were evaluated. PC-mNGS achieved higher pathogen detection rates (66.67% vs 57.10%, P = 0.000198) than mNGS, particularly in blood (66.67% vs 47.52%, P = 2.5 × 10⁻⁵). PC-mNGS detected more bacteria (19 species exclusive) and fungi (11 species exclusive) than mNGS. Viruses showed comparable detection. BALF and CSF exhibited high overall agreement (OPA: 96.30% and 88%, respectively), while blood had lower concordance (NPA: 54.05%, OPA: 70.92%). A total of 60.55% of BALF-positive samples (PC-mNGS) had co-detected pathogens in blood. Gram-negative bacteria (e.g., Klebsiella pneumoniae) and fungi (e.g., Candida albicans) showed higher blood co-detection rates than viruses. In this study, PC-mNGS detected more pathogens and showed a higher positivity rate than mNGS in blood samples. BALF sequencing data, particularly bacterial reads per million (RPM), may predict bloodstream co-detection, aiding in sepsis management. However, clinical validation and integration with traditional diagnostics are needed to confirm utility. This study highlights PC-mNGS as a promising tool for complex infections but underscores the need for rigorous multi-context validation.IMPORTANCEAccurate and rapid identification of pathogens is critical for effective treatment of severe infectious diseases, such as sepsis. This study demonstrates that probe-capture metagenomic sequencing (PC-mNGS) detected more pathogens in blood samples compared to conventional metagenomic sequencing, especially for bacterial and fungal infections. By analyzing paired lung and blood samples, we show that high pathogen levels in lung fluid may predict bloodstream infection, offering a potential early warning for clinicians. These findings support the use of PC-mNGS as a more sensitive diagnostic tool, which could lead to faster, more targeted therapies and better outcomes for patients with complex infections.

Humans

A hybrid and cost-efficient barcoding strategy for full-length 16S rRNA gene nanopore sequencing of environmental samples.

BACKGROUND: Accurate species-level identification of bacteria in complex environmental samples is essential for applications in biotechnology, ecological monitoring, and clinical diagnostics. Short-read platforms such as Illumina frequently truncate the 16S rRNA gene, limiting taxonomic resolution. In this work, we applied Oxford Nanopore Technology (ONT) long-read sequencing to full-length 16S rRNA amplicon in samples from natural soil amended with lignocellulosic biomass and a simplified microbial community derived from cultures grown on selective and differential carboxymethyl cellulose (CMC)-based substrates, with the aim to evaluate the difference in performance between a real, complex community and a less complex system. To reduce consumable costs, we substituted the standard ONT Barcoding kits with an in-house hybrid barcoding workflow. Specifically, PacBio PCR-based barcoding protocol was used for sample indexing, followed by library preparation using the ONT Ligation Sequencing Kit. This simplified approach retained compatibility with MinION and Flongle flow cells and supported accurate downstream demultiplexing while lowering barcode costs substantially. Additionally, a new bioinformatic workflow tailored to ONT data was implemented. RESULTS: Overall, the hybrid protocol significantly reduced per-sample barcoding costs while preserving high sequencing quality and throughput. The sequencing run yielded over 5 Gb of quality-filtered data (Q-score ≥ 10). Furthermore, the new bioinformatic workflow allowed taxonomic assignment at the species level for 49.38% of annotated taxa, compared to just 4.59% using Illumina NovaSeq sequencing of the V3-V4 region. ONT also recovered 2.3 times more genera and 1.3 times more families. Although 16S rRNA gene sequencing often cannot distinguish between closely related species, particularly within taxonomically complex groups, in this work, full-length reads substantially improved both taxonomic resolution and database matching. CONCLUSIONS: These results show that full-length 16S rRNA sequencing with ONT, paired with a low-cost barcoding strategy, enhanced taxonomic resolution compared to short-read workflows. This approach also offers a scalable and cost-effective option for high-resolution microbiome profiling in research and applied settings.

RNA, Ribosomal, 16S

Exploring the use of machine and deep learning in genome-wide association studies: a comprehensive review.

The advent of high-throughput sequencing technologies has generated increasingly large and complex genomic datasets, necessitating analytical approaches capable of capturing high-dimensional and potentially nonlinear genetic interactions. This situation has significantly impacted the entire field of Genome-Wide Association Study (GWAS), whose primary goal is the identification of genomic traits and variants that are statistically associated with the risk of a disease. However, traditional GWAS methods may show reduced performance when applied to highly polygenic and nonlinear genetic architectures. Computational strategies from Artificial Intelligence (AI) and, in particular, from machine- and deep-learning may provide a powerful tool to overcome such limitations, especially by capturing nonlinear interactions and complex hidden regularities in large-scale data, which traditional GWAS approaches might overlook. To date, only a few approaches have been introduced and systematically assessed. In this review, we describe the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS. Particular attention will be devoted to key issues, such as the interpretability of methods and results, and the curse of dimensionality. More specifically, the review presents 30 methods designed to leverage AI in GWAS, as well as presenting a comprehensive set of evaluation metrics for their performance, also providing references to the most frequently used databases, and biobanks. Overall, this work may serve as a starting point for both dry- and wet-lab researchers, aiming to extract deeper insights from genomic data by moving beyond traditional linear additive assumptions, and leveraging large-scale datasets through AI-driven approaches.

Artificial intelligence

[Applications and Challenges of Deep Learning in Human Genome Research].

In recent years, the advent of high-throughput omics technologies has fueled an explosive growth in human genomic data. Uncovering the latent functions within this vast data has become a significant challenge in functional genomics research. While traditional statistical methods have proved successful for analyzing smaller-scale datasets in the past, they exhibit clear limitations in analytical efficiency and integrating multi-dimensional data, struggling to meet the escalating demands of contemporary genomic analysis. The introduction of deep learning (DL) technologies offers a novel paradigm for this field. This review systematically examines the advances in applying deep learning to human genomics research. Studies demonstrate that when ample labeled data is available, discriminative DL computational methods-such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs)-achieve high accuracy and efficiency in genomic variant discovery tasks. Furthermore, generative DL methods, particularly Large Language Models (LLMs) leveraging self-supervised pre-training strategies, effectively integrate complex genomic information and exhibit superior performance in functional genomic sequence annotation and gene regulation studies. This review also explores the application of LLMs in multi-omics data integration and prediction. Looking ahead, the continued accumulation of long-read sequencing and high-dimensional data is expected to enable DL technologies to integrate increasingly complex and heterogeneous genomic information, playing an increasingly crucial role in human genomics research.

Deep Learning

Tumor Mutational Landscape and Its Correlation With Histopathological Characteristics in Breast Cancer.

BACKGROUND/AIM: In breast cancer, knowledge of the associations between clinicopathologic characteristics, genetic changes, and subtype-specific patterns is expanding. This study investigated how pathological and clinical variables affect the actionability of Next Generation Sequencing (NGS)-based tumor molecular data. MATERIALS AND METHODS: 227 breast cancer patients referred to Genekor's laboratory for tumor molecular profile analysis were included in the study. Pathology records were used to assess critical clinicopathological features, including HER2, ER, PR, Ki67, grade, metastatic site, and age. A 1021-gene NGS-based multigene panel was utilized to assess tumor biology alongside tumor mutational burden (TMB) and microsatellite instability (MSI). RESULTS: Comprehensive genomic profiling revealed that 95.6% of the patients harbored at least one oncogenic or likely oncogenic alteration, highlighting the high diagnostic yield of NGS-based testing. Distinct subtype-specific patterns were observed: HR+/HER2- tumors were enriched for PIK3CA and ESR1 gene alterations, whereas triple-negative breast cancer (TNBC) was dominated by TP53 alterations. Clinically actionable alterations were most common in HR+/HER2- tumors (~60% on-label), whereas TNBC more often harbored off-label or trial-associated targets. The inclusion of tumor-agnostic biomarkers (TMB/MSI) increased on-label actionability up to 64.5% in HR+/HER2- tumors, primarily driven by TMB-high cases. Median TMB values were low, and age was the only independent predictor. Furthermore, the presence of actionable alterations was significantly higher in metastatic tumors, and TP53 alterations were associated with aggressive tumor characteristics. CONCLUSION: Comprehensive NGS-based genomic profiling identifies clinically actionable alterations in over half of breast cancer patients, with substantial variability across molecular subtypes. The HR+/HER2- subtype demonstrates the highest prevalence of on-label actionable biomarkers. These findings support the routine implementation of comprehensive genomic profiling, especially in metastatic HER2-negative breast cancer, to guide precision oncology strategies and enable enrollment in biomarker-driven clinical trials.

Humans

Lawsonella clevelandensis: a normal flora that bites deep.

Since its formal description in 2016, Lawsonella clevelandensis-a strictly anaerobic, partially acid-fast bacterium-has been increasingly recognized as a cause of deep-seated abscesses, yet its fastidious nature and absence from routine diagnostic databases contribute to significant underdiagnosis. This narrative review synthesizes current knowledge on its microbiology, expanding clinical spectrum, diagnostic strategies, and treatment, based on a literature search of PubMed and Web of Science up to April 2026. Analysis of 27 documented publications, comprising 18 clinical cases, reveals a potential association with host risk factors including diabetes, immunosuppression, and prior surgical procedures, alongside a notable predilection for fat-rich tissues such as the breast and abdomen. While metagenomic next-generation sequencing and 16S rRNA gene amplification have become indispensable for definitive identification, antimicrobial susceptibility data-derived primarily from a single strain-demonstrate uniformly low minimum inhibitory concentrations for penicillins, carbapenems, clindamycin, and metronidazole, with no acquired resistance genes identified by whole-genome sequencing. However, the absence of established clinical breakpoints and limited tested isolates precludes definitive conclusions about universal susceptibility. Clinicians should maintain a high index of suspicion for L. clevelandensis in culture-negative deep abscesses, particularly those with acid-fast rods, as prompt diagnosis and empirical therapy with β-lactam/β-lactamase inhibitors or carbapenems appear reasonable based on current in vitro and clinical evidence, though further susceptibility surveillance is essential.

Humans