PubMed HealthSearch

SEARCH · PubMed Health

Results for “High-throughput sequencing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans

DirectASRM: uncovering allele-specific post-transcriptional RNA modifications through direct RNA sequencing.

SUMMARY: We developed DirectASRM, a comprehensive database for the systematic identification, integration, and annotation of allele-specific RNA modifications (ASRMs) from direct RNA sequencing data. DirectASRM enables single-base, transcript-level detection of ASRMs across multiple RNA modification types, diverse organisms and condition-specific contexts. The database further evaluates the confidence of each ASRM-SNP pair association within isoform context by jointly considering statistical evidence of allelic modification imbalance and independent support from external next-generation sequencing (NGS) - based RNA modification resources. DirectASRM also provides extensive functional annotations for ASRMs and their associated variants, including intra-sample transcript-level allele-specific expression (ASE) and allele-specific splicing, as well as additional post-transcriptional regulatory features such as miRNA binding, circRNA, RNA-protein interactions, and disease relevance. Overall, DirectASRM serves as a comprehensive resource that supports systematic investigation of the potential functional impact of genetic variants in epitranscriptomic regulation. AVAILABILITY AND IMPLEMENTATION: DirectASRM database is freely accessible at http://modinfor.com/DirectASRM/. DirectASRM pipeline is available at GitHub (https://github.com/jiayin1101/DirectASRM_pipeline) and Zenodo (DOI: https://doi.org/10.5281/zenodo.19876077).

Alleles

Analysis of deep-resequencing data of 984 soybean accessions reveals structural variations underlying agronomic traits.

Genomic structural variants (SVs) are major sources of genetic variation and have profound impacts on phenotypic traits. However, their functional effects remain largely unexplored in soybean. Here, we resequence 940 soybean accessions. Together with 44 publicly available datasets, we identify 602,281 SVs. Using a graph-based genome, we detect an additional 58,760 presence/absence variations (PAVs) that broadly affect gene expression. Population genomic analyses reveal that SVs serve as a core driving force for soybean domestication and improvement. Integrating SVs with QTLs for oil and protein content, and performing GWAS on 27 traits, we identify key functional SVs. These include transposable element insertions altering seed coat color, multiple insertions within a cytochrome P450 gene modifying flower and hypocotyl color, and a GmMATE1 deletion enhancing seed size. Together, our study establishes a comprehensive SV map of soybean, offering a valuable resource for dissecting the genetic basis of complex traits to accelerate molecular breeding.

Glycine max

Optimizing GRIDSS for clinical use: A targeted NGS filtering strategy for germline structural variant detection.

Detecting intermediate-sized structural variants (SVs) remains challenging in diagnostics, as tools for single-nucleotide and copy-number variants, particularly read-depth-based methods, are often insufficient. GRIDSS addresses this gap by integrating paired-end mapping, split-read analysis, and assembly-based approaches. However, its use in targeted sequencing and diagnostic workflows remains complex. NGS panel data from 9726 patients with suspected hereditary cancer were analyzed using GRIDSS. A filtering strategy was developed to prioritize clinically relevant germline SVs. Multiple parameter settings were tested to optimize performance. The initial dataset of 1,307,592 variants was reduced to 89 candidates after applying the selected filtering strategy. Of these, 24 had been previously detected by routine callers and were not further analyzed. Among the remaining 65, 13 were considered likely true positives after visual inspection using IGV. Experimental validation was performed by Sanger/Nanopore long-read sequencing for these variants, all of which were confirmed. Eight were classified as (likely) pathogenic, including two frameshift duplications in MSH6, one splicing variant in BARD1, and five mobile element insertions in APC, BRCA2, and PALB2. Altogether, GRIDSS implementation increased diagnostic yield while maintaining feasibility for diagnostic workflows. Comprehensive workflow scheme for germline structural variant detection and results in our diagnostic setting.

Humans

Lynch syndrome-associated urothelial carcinoma: clinical and molecular findings from a single-institution cohort.

Lynch syndrome-associated urothelial carcinoma (LS-UC) is a rare and undercharacterized clinical entity. While FGFR3 alterations are well described in sporadic urothelial carcinoma, their prevalence and clinical implications in LS-UC remain unclear. We aimed to provide a comprehensive clinical and molecular characterization of LS-UC. We conducted a retrospective single-center study including patients with Lynch syndrome (LS) and histologically confirmed urothelial carcinoma (UC). Clinical, pathological, treatment, and follow-up data were collected. Targeted next-generation sequencing was performed on available tumor samples to assess genomic alterations, with particular attention to FGFR3 mutations. A total of 27 patients with LS-UC were identified, with a predominance of upper urinary tract involvement (70%). Most tumors were diagnosed at an early stage and initially managed with local treatment. During a median follow-up of 92 months, 48% of patients experienced recurrence, with a median time to recurrence of 37 months. Recurrences were predominantly local and were mainly managed with additional surgical or intravesical treatments. No deaths were attributable to UC at last follow-up. Molecular analysis was feasible in 9 cases. FGFR3 mutations were detected in 67% of evaluable samples, with the recurrent p.Arg248Cys hotspot identified in 55% of cases. Additional alterations involved TP53, SWI/SNF complex genes, and PIK3CA, which co-occurred with FGFR3 p.Arg248Cys. No gene fusions were identified. This study expands the limited molecular and clinical evidence on Lynch syndrome-associated urothelial carcinoma. Beyond confirming the recurrent role of FGFR3 (notably p.Arg248Cys), our comprehensive multigene profiling enriches the current genomic knowledge for this rare population. Multi-center collaborative efforts remain essential to aggregate larger datasets and ultimately guide personalized patient management.

Humans

Bilateral Conversion Risk in Unilateral Retinoblastoma Using Age and Genetic Testing.

IMPORTANCE: Metachronous bilateral conversion in initially unilateral retinoblastoma is uncommon but clinically consequential, potentially requiring intensified treatment and carrying worse prognosis. Clarifying how age at diagnosis refines genetic-risk stratification could enable safer, more efficient surveillance protocols. OBJECTIVE: To estimate the incidence and timing of metachronous bilateral conversion in unilateral retinoblastoma and assess whether age at diagnosis and RB1 testing are associated with bilateral conversion risk. DESIGN, SETTING AND PARTICIPANTS: This was a retrospective cohort study at a tertiary center in Shanghai, China, including 1108 consecutive children with initially unilateral retinoblastoma diagnosed from July 2010 to October 2024 (after exclusions for short follow-up [n = 139], missing data [n = 53], or synchronous bilateral disease [n = 10]). The median (IQR) follow-up was 43.4 (24.2-67.6) months. EXPOSURES: Age at diagnosis and RB1 genetic status/subtypes assessed by next-generation sequencing and multiplex ligation-dependent probe amplification, including penetrance class (high vs low) and mosaic vs germline categorization. MAIN OUTCOMES AND MEASURES: Time to metachronous bilateral conversion; cumulative incidence functions with death as a competing risk; spatial distribution of fellow-eye tumors. RESULTS: Among 1108 patients (median [IQR] age at diagnosis, 22.2 [12.0-31.4] months; 591 [53.3%] male), 24 (2.2%) developed metachronous bilateral disease. At 24 months, cumulative incidence was 2.2% (95% CI, 1.3-3.1) overall. By genetic status, the 24-month cumulative incidence was 24.8% (95% CI, 13.8-35.9) in RB1 variant-positive vs 1.6% (95% CI, 0.0-3.1) in RB1 variant-negative patients. Among RB1 variant-positive patients, risk clustered among those diagnosed before 9 months, whereas no conversions were observed among those diagnosed at older than 9 months. Four RB1 variant-negative patients who were initially diagnosed at notably late ages (20.9, 42.7, 79.6, and 118 months) subsequently converted; these cases likely represent undetected low-level mosaicism, somatic variants below detection thresholds, or rare genomic events not captured by standard sequencing panels. Fellow-eye tumors did not involve macula and showed a nasal-predominant distribution. CONCLUSIONS AND RELEVANCE: The findings in this study suggest that age at diagnosis may refine genetic risk stratification for metachronous bilateral conversion. RB1 variant-positive patients diagnosed at 9 months or later represent a very low-risk subgroup that may warrant surveillance deescalation, while rare late conversions in RB1 variant-negative patients necessitate continued long-term monitoring.

Humans

Obtaining a Diagnostic Yield via Scan findings prior to the introduction of SEquencing retrospectivelY (ODYSSEY): a cohort study.

OBJECTIVE: To determine the retrospective yield of prenatal exome sequencing (PES) by establishing the proportion of children with a postnatal monogenic diagnosis that could have been diagnosed prenatally if PES had been available. METHODS: The study cohort comprised a sample of children in Northern Ireland, born between January 2010 and January 2018 (predating routine availability of PES), who received a monogenic diagnosis postnatally via next generation sequencing as part of either of two UK-wide studies (the 100 000 Genomes Project (2015-2018) or the Deciphering Developmental Disorders study (2011-2015)). Clinical data were collected retrospectively and correlated with the current UK National Health Service PES protocol, including the phenotypic eligibility criteria for PES and the associated fetal anomalies gene panel. Cases were considered retrospective diagnoses if the fetal phenotype would have been eligible for PES and the diagnostic gene was included on the test panel, meaning prenatal diagnosis in this current era could have been feasible. RESULTS: Of 101 children, 17.8% (95% CI, 10.3-25.3%) had both an eligible fetal structural anomaly (FSA) (i.e. high-risk FSA) and a diagnostic gene on the associated test panel, meaning that they could have been diagnosed prenatally in the current clinical landscape. The median length of the diagnostic odyssey for this subgroup of children was 3.7 years (1354 (range, 822-2450) days). Moreover, 58.4% (n = 59) of cases had no anomalies detected prenatally and 19.8% (n = 20) had a FSA that would not meet the eligibility criteria for PES (low-risk FSA). Although these cases would have been ineligible for PES under the current clinical pathway, 89.9% (n = 71/79) were affected by severe or profound syndromes. Postnatally, the most common functional anomalies were neurodevelopmental delay/intellectual disability and/or behavioral abnormality, which were observed in 80.2% (n = 81) of the included children. However, 80.2% (n = 65/81) of these affected children did not present with fetal anomalies eligible for PES. CONCLUSIONS: Almost one-fifth of children with a monogenic condition included in this study could have received a diagnosis via modern PES, avoiding a diagnostic odyssey lasting almost 4 years. However, despite having a monogenic condition, over half of the children did not present with any structural anomalies in utero. This demonstrates the degree to which fetal imaging is limited in its ability to reassure parents of the absence of a fetal genetic syndrome. © 2026 The Author(s). Ultrasound in Obstetrics & Gynecology published by John Wiley & Sons Ltd on behalf of International Society of Ultrasound in Obstetrics and Gynecology.

Humans

Characterisation of Bordetella pertussis virulence and macrolide resistance in Australia by targeted culture-independent sequencing: a genomic epidemiology study.

BACKGROUND: Bordetella pertussis continues to circulate globally despite widespread vaccination, with a notable epidemic in 2024. Its resurgence is confounded by the emergence of pertactin-deficient, macrolide-resistant B pertussis strains in Asia and Europe, which are under-recognised by conventional diagnostics. We aimed to apply targeted culture-independent next-generation sequencing (tNGS) of respiratory specimens to improve global B pertussis diagnostic capability and genomic surveillance. METHODS: We did a nationwide genomic epidemiology study of B pertussis RT-PCR-positive respiratory specimens that were retrospectively and prospectively collected by diagnostic and public health laboratories in six of seven states and territories of Australia. Specimens underwent tNGS and macrolide-resistant B pertussis-specific PCR, and an opportunistic subset from New South Wales and Queensland were cultured for confirmatory susceptibility testing and whole-genome sequencing. Sequencing data were analysed for genome recovery, virulence profiles, and macrolide resistance mutations, and were compared with international macrolide-resistant B pertussis genomes and ancestral Australian genomes. The performance of the tNGS approach was assessed with logistic regression relative to RT-PCR cycle threshold values, and sensitivity and specificity values were calculated. FINDINGS: 255 respiratory specimens positive for B pertussis were included in the study. 64 (25%) were retrospectively collected between Jan 12, 2012, and Dec 31, 2023, and 191 (75%) were prospectively collected between Jan 1 and Oct 28, 2024. Of these 255 specimens, 148 (58%) yielded near-complete B pertussis genomes through tNGS. Seven co-circulating lineages of B pertussis were documented, including two associated with macrolide-resistance. Eight epidemiologically unrelated and geographically dispersed cases of macrolide-resistant B pertussis with a 23S rRNA 2037A→G mutation were identified by tNGS and confirmed by whole-genome sequencing. Three of these were further validated by phenotypic testing. The estimated prevalence of macrolide resistance among Australian cases positive for B pertussis was 4% (eight of 188). INTERPRETATION: tNGS can recover near-complete B pertussis genomes directly from clinical specimens, enabling identification of macrolide resistance mutations and high-resolution phylogenetic analysis. These findings show that tNGS complements PCR-based surveillance by providing genome-wide assessment of resistance, virulence, and genomic diversity in a single workflow. FUNDING: NSW Health Prevention Research Support Program.

Macrolides

Probability of Mitochondrial DNA heteroplasmy in different tissues from European populations.

Mitochondrial DNA (mtDNA) heteroplasmy complicates genetic analyses due to its variability across individuals and tissues. We analyzed over 400 Spanish blood samples and integrated published Massively Parallel Sequencing (MPS) data from ten additional European tissues. Heteroplasmy was tissue-specific, with skeletal muscle, kidney, and liver showing the highest levels, while the intestines, skin, and cerebellum had the lowest. Blood uniquely displayed more heteroplasmies in coding than non-coding regions. Several conserved positions not previously described as hotspots showed high frequencies. These results establish the first comprehensive tissue-specific heteroplasmic profile of the complete mitochondrial genome in a European population, improving the interpretation of mtDNA variation in forensic and biomedical contexts.

Humans

Clinical Utility of Next-Generation Sequencing in Tumors Diagnosed as Lung Squamous Cell Carcinoma: Real-World Data of Diagnostic and Therapeutic Implications.

Lung squamous cell carcinoma (LUSC) is the second most common subtype of non-small cell lung carcinoma (NSCLC), typically associated with a poor prognosis. Unlike lung adenocarcinoma, the application of next-generation sequencing (NGS) in LUSC has lagged because of the long-standing perception of low therapeutic yield, primarily based on highly selected, resected cohorts. We sought to determine the real-world clinical utility of NGS in LUSC. We analyzed an institutional cohort of 576 tumors initially diagnosed as LUSC that underwent NGS profiling. We defined "clinical yield" as either diagnostic reclassification or the identification of a targetable mitogenic alteration. Twenty cases (3.5%) were reclassified, including rediagnosis to cutaneous squamous cell carcinoma, transformed adenocarcinoma (post targeted therapy), and rare entities such as nuclear protein of the testis-rearranged carcinoma and lymphoepithelial carcinoma. Primary mitogenic drivers were identified in 83 cases (14.4% of the total cohort), of which 43 (7.5% of the total cohort) harbored alterations with currently Food and Drug Administration-approved therapies for NSCLC (including KRAS, EGFR, MET, ALK, and ROS1). Overall clinical yield-defined as the sum of diagnostic reclassifications and identification of NSCLC-specific targetable alterations-was 11.0% (63/576). Univariate and multivariate analysis demonstrated that never or light smoking history was the strongest independent predictor of clinical yield, with 57.3% of tumors in this subset being reclassified or harboring a strong driver. Our findings demonstrate that NGS provides significant diagnostic and therapeutic value in a real-world LUSC cohort, challenging the historical premise of low yield. Although clinicodemographic features can help prioritize testing in resource-limited settings, the identification of targetable drivers across all smoking groups supports the universal application of comprehensive NGS for all patients diagnosed with LUSC.

Humans

Identification of autosomal and sex chromosome aneuploidies using next generation sequencing.

MOTIVATION: Chromosomal abnormalities, referred to as aneuploidies, occur in approximately 0.3% of live births. While the majority of aneuploidies in humans are incompatible with life, well-characterized exceptions include Down syndrome (47,+21), Patau syndrome (47,+13), Edwards syndrome (47,+18), Turner syndrome (45,X0), Klinefelter syndrome (47,XXY), and triple X syndrome (47,XXX). These chromosomal alterations disrupt gene expression and cellular function, leading to genetic and developmental disorders. With the increasing adoption of next generation sequencing (NGS) in clinical diagnostics, this study aims to explore the potential use of NGS for aneuploidies detection. RESULTS: Using data derived from clinical exomes (CES) and whole exomes (WES) sequencing we have been able to detect autosomal as well as sex chromosome aneuploidies with high specificity. Moreover, we have also been able to identify mosaic aneuploidies proving the high sensibility of this methodological approach. Thus, we present NGS as a cost-effective first line approach to detect chromosomal aneuploidies in routine diagnostic practice. AVAILABILITY AND IMPLEMENTATION: Scripts are available at https://github.com/B-R-I-D-G-E/AneuploidiesStudies.

Humans

Assessing the readiness of Oxford Nanopore sequencing for clinical genomics applications.

Long-read sequencing (LRS) technologies, namely, Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio), have emerged as promising solutions to overcome the limitations of short-read sequencing (SRS). Nevertheless, the still higher sequencing error rates compared with SRS, need for customized pipelines, rapidly updating software, and incipient scalability continue to present challenges for adopting ONT in standard clinical practice. Here we assess the performance of ONT (R9 and R10 chemistries) in comparison to Illumina and MGI across 17 well-characterized reference samples with 11 clinical variants representing nine different genetic diseases. To enable this, we have implemented a production-ready pipeline including SNV, indel, STR, SV, and CNV detection, alongside reporting key summary metrics to ensure high-quality data at the production sequencing level. Our results show high accuracy of ONT across SNVs (F-score 0.978-0.983) and SVs (F-score = 0.75) but still weaknesses across indels (F-score 0.659-0.758). However, we highlight that ONT accurately detected all four pathogenic indels as well as the performance improvement in exons and with the newer R10 chemistry. We further demonstrated the importance of long reads to detect clinically impactful variants such as a FMR1 pathogenic expansion, often misclassified by SRS as being in the premutation range. Our multiplatform analysis and Sanger validation uncovered a 1 bp error in the Coriell annotation for a cystic fibrosis-causing indel in GM07829. This work underscores the growing readiness of ONT for clinical applications, highlighting both its advancements and its potential for broader adoption in clinical genomics and large-scale operations.

Humans

Comparative evaluation of probe-capture and conventional metagenomic sequencing across multiple clinical sample types, with analysis of paired bronchoalveolar lavage fluid and blood samples.

Conventional metagenomic next-generation sequencing (mNGS) suffers from host nucleic acid interference and poor performance in low-biomass samples. Probe-capture metagenomic sequencing (PC-mNGS), which enriches microbial targets via hybridization probes, shows superior sensitivity but lacks systematic multi-sample evaluations. This study compared PC-mNGS and mNGS across diverse clinical specimens (bronchoalveolar lavage fluid [BALF], blood, cerebrospinal fluid [CSF]) and assessed the clinical utility of pathogen co-detection in paired BALF-blood samples from sepsis patients. A total of 282 samples (81 BALF, 141 blood, 25 CSF, 35 others) sequenced by both PC-mNGS and mNGS were analyzed. Additionally, 621 paired BALF-blood samples from sepsis patients with pulmonary infections were evaluated. PC-mNGS achieved higher pathogen detection rates (66.67% vs 57.10%, P = 0.000198) than mNGS, particularly in blood (66.67% vs 47.52%, P = 2.5 × 10⁻⁵). PC-mNGS detected more bacteria (19 species exclusive) and fungi (11 species exclusive) than mNGS. Viruses showed comparable detection. BALF and CSF exhibited high overall agreement (OPA: 96.30% and 88%, respectively), while blood had lower concordance (NPA: 54.05%, OPA: 70.92%). A total of 60.55% of BALF-positive samples (PC-mNGS) had co-detected pathogens in blood. Gram-negative bacteria (e.g., Klebsiella pneumoniae) and fungi (e.g., Candida albicans) showed higher blood co-detection rates than viruses. In this study, PC-mNGS detected more pathogens and showed a higher positivity rate than mNGS in blood samples. BALF sequencing data, particularly bacterial reads per million (RPM), may predict bloodstream co-detection, aiding in sepsis management. However, clinical validation and integration with traditional diagnostics are needed to confirm utility. This study highlights PC-mNGS as a promising tool for complex infections but underscores the need for rigorous multi-context validation.IMPORTANCEAccurate and rapid identification of pathogens is critical for effective treatment of severe infectious diseases, such as sepsis. This study demonstrates that probe-capture metagenomic sequencing (PC-mNGS) detected more pathogens in blood samples compared to conventional metagenomic sequencing, especially for bacterial and fungal infections. By analyzing paired lung and blood samples, we show that high pathogen levels in lung fluid may predict bloodstream infection, offering a potential early warning for clinicians. These findings support the use of PC-mNGS as a more sensitive diagnostic tool, which could lead to faster, more targeted therapies and better outcomes for patients with complex infections.

Humans

A hybrid and cost-efficient barcoding strategy for full-length 16S rRNA gene nanopore sequencing of environmental samples.

BACKGROUND: Accurate species-level identification of bacteria in complex environmental samples is essential for applications in biotechnology, ecological monitoring, and clinical diagnostics. Short-read platforms such as Illumina frequently truncate the 16S rRNA gene, limiting taxonomic resolution. In this work, we applied Oxford Nanopore Technology (ONT) long-read sequencing to full-length 16S rRNA amplicon in samples from natural soil amended with lignocellulosic biomass and a simplified microbial community derived from cultures grown on selective and differential carboxymethyl cellulose (CMC)-based substrates, with the aim to evaluate the difference in performance between a real, complex community and a less complex system. To reduce consumable costs, we substituted the standard ONT Barcoding kits with an in-house hybrid barcoding workflow. Specifically, PacBio PCR-based barcoding protocol was used for sample indexing, followed by library preparation using the ONT Ligation Sequencing Kit. This simplified approach retained compatibility with MinION and Flongle flow cells and supported accurate downstream demultiplexing while lowering barcode costs substantially. Additionally, a new bioinformatic workflow tailored to ONT data was implemented. RESULTS: Overall, the hybrid protocol significantly reduced per-sample barcoding costs while preserving high sequencing quality and throughput. The sequencing run yielded over 5 Gb of quality-filtered data (Q-score ≥ 10). Furthermore, the new bioinformatic workflow allowed taxonomic assignment at the species level for 49.38% of annotated taxa, compared to just 4.59% using Illumina NovaSeq sequencing of the V3-V4 region. ONT also recovered 2.3 times more genera and 1.3 times more families. Although 16S rRNA gene sequencing often cannot distinguish between closely related species, particularly within taxonomically complex groups, in this work, full-length reads substantially improved both taxonomic resolution and database matching. CONCLUSIONS: These results show that full-length 16S rRNA sequencing with ONT, paired with a low-cost barcoding strategy, enhanced taxonomic resolution compared to short-read workflows. This approach also offers a scalable and cost-effective option for high-resolution microbiome profiling in research and applied settings.

RNA, Ribosomal, 16S

Exploring the use of machine and deep learning in genome-wide association studies: a comprehensive review.

The advent of high-throughput sequencing technologies has generated increasingly large and complex genomic datasets, necessitating analytical approaches capable of capturing high-dimensional and potentially nonlinear genetic interactions. This situation has significantly impacted the entire field of Genome-Wide Association Study (GWAS), whose primary goal is the identification of genomic traits and variants that are statistically associated with the risk of a disease. However, traditional GWAS methods may show reduced performance when applied to highly polygenic and nonlinear genetic architectures. Computational strategies from Artificial Intelligence (AI) and, in particular, from machine- and deep-learning may provide a powerful tool to overcome such limitations, especially by capturing nonlinear interactions and complex hidden regularities in large-scale data, which traditional GWAS approaches might overlook. To date, only a few approaches have been introduced and systematically assessed. In this review, we describe the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS. Particular attention will be devoted to key issues, such as the interpretability of methods and results, and the curse of dimensionality. More specifically, the review presents 30 methods designed to leverage AI in GWAS, as well as presenting a comprehensive set of evaluation metrics for their performance, also providing references to the most frequently used databases, and biobanks. Overall, this work may serve as a starting point for both dry- and wet-lab researchers, aiming to extract deeper insights from genomic data by moving beyond traditional linear additive assumptions, and leveraging large-scale datasets through AI-driven approaches.

Artificial intelligence

[Applications and Challenges of Deep Learning in Human Genome Research].

In recent years, the advent of high-throughput omics technologies has fueled an explosive growth in human genomic data. Uncovering the latent functions within this vast data has become a significant challenge in functional genomics research. While traditional statistical methods have proved successful for analyzing smaller-scale datasets in the past, they exhibit clear limitations in analytical efficiency and integrating multi-dimensional data, struggling to meet the escalating demands of contemporary genomic analysis. The introduction of deep learning (DL) technologies offers a novel paradigm for this field. This review systematically examines the advances in applying deep learning to human genomics research. Studies demonstrate that when ample labeled data is available, discriminative DL computational methods-such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs)-achieve high accuracy and efficiency in genomic variant discovery tasks. Furthermore, generative DL methods, particularly Large Language Models (LLMs) leveraging self-supervised pre-training strategies, effectively integrate complex genomic information and exhibit superior performance in functional genomic sequence annotation and gene regulation studies. This review also explores the application of LLMs in multi-omics data integration and prediction. Looking ahead, the continued accumulation of long-read sequencing and high-dimensional data is expected to enable DL technologies to integrate increasingly complex and heterogeneous genomic information, playing an increasingly crucial role in human genomics research.

Deep Learning

Tumor Mutational Landscape and Its Correlation With Histopathological Characteristics in Breast Cancer.

BACKGROUND/AIM: In breast cancer, knowledge of the associations between clinicopathologic characteristics, genetic changes, and subtype-specific patterns is expanding. This study investigated how pathological and clinical variables affect the actionability of Next Generation Sequencing (NGS)-based tumor molecular data. MATERIALS AND METHODS: 227 breast cancer patients referred to Genekor's laboratory for tumor molecular profile analysis were included in the study. Pathology records were used to assess critical clinicopathological features, including HER2, ER, PR, Ki67, grade, metastatic site, and age. A 1021-gene NGS-based multigene panel was utilized to assess tumor biology alongside tumor mutational burden (TMB) and microsatellite instability (MSI). RESULTS: Comprehensive genomic profiling revealed that 95.6% of the patients harbored at least one oncogenic or likely oncogenic alteration, highlighting the high diagnostic yield of NGS-based testing. Distinct subtype-specific patterns were observed: HR+/HER2- tumors were enriched for PIK3CA and ESR1 gene alterations, whereas triple-negative breast cancer (TNBC) was dominated by TP53 alterations. Clinically actionable alterations were most common in HR+/HER2- tumors (~60% on-label), whereas TNBC more often harbored off-label or trial-associated targets. The inclusion of tumor-agnostic biomarkers (TMB/MSI) increased on-label actionability up to 64.5% in HR+/HER2- tumors, primarily driven by TMB-high cases. Median TMB values were low, and age was the only independent predictor. Furthermore, the presence of actionable alterations was significantly higher in metastatic tumors, and TP53 alterations were associated with aggressive tumor characteristics. CONCLUSION: Comprehensive NGS-based genomic profiling identifies clinically actionable alterations in over half of breast cancer patients, with substantial variability across molecular subtypes. The HR+/HER2- subtype demonstrates the highest prevalence of on-label actionable biomarkers. These findings support the routine implementation of comprehensive genomic profiling, especially in metastatic HER2-negative breast cancer, to guide precision oncology strategies and enable enrollment in biomarker-driven clinical trials.

Humans

Lawsonella clevelandensis: a normal flora that bites deep.

Since its formal description in 2016, Lawsonella clevelandensis-a strictly anaerobic, partially acid-fast bacterium-has been increasingly recognized as a cause of deep-seated abscesses, yet its fastidious nature and absence from routine diagnostic databases contribute to significant underdiagnosis. This narrative review synthesizes current knowledge on its microbiology, expanding clinical spectrum, diagnostic strategies, and treatment, based on a literature search of PubMed and Web of Science up to April 2026. Analysis of 27 documented publications, comprising 18 clinical cases, reveals a potential association with host risk factors including diabetes, immunosuppression, and prior surgical procedures, alongside a notable predilection for fat-rich tissues such as the breast and abdomen. While metagenomic next-generation sequencing and 16S rRNA gene amplification have become indispensable for definitive identification, antimicrobial susceptibility data-derived primarily from a single strain-demonstrate uniformly low minimum inhibitory concentrations for penicillins, carbapenems, clindamycin, and metronidazole, with no acquired resistance genes identified by whole-genome sequencing. However, the absence of established clinical breakpoints and limited tested isolates precludes definitive conclusions about universal susceptibility. Clinicians should maintain a high index of suspicion for L. clevelandensis in culture-negative deep abscesses, particularly those with acid-fast rods, as prompt diagnosis and empirical therapy with β-lactam/β-lactamase inhibitors or carbapenems appear reasonable based on current in vitro and clinical evidence, though further susceptibility surveillance is essential.

Humans