PubMed HealthSearch

SEARCH · PubMed Health

Results for “High-throughput sequencing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

NextLongIso: a comprehensive Nextflow pipeline for multi-dimensional long-read RNA-seq analysis.

SUMMARY: Long-read RNA sequencing technologies, including Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT), enable direct characterization of full-length transcripts and transcriptome complexity. However, analysis of long-read RNA-seq data remains fragmented across multiple tools, limiting the ability to obtain a unified view of transcript structure, expression, and regulatory variation in long-read transcriptomes. We present NextLongIso, a scalable and reproducible Nextflow pipeline that enables coordinated analysis of multiple layers of transcript regulation. Rather than focusing solely on transcript reconstruction, NextLongIso integrates transcript discovery with downstream regulatory analyses to jointly characterize alternative splicing, isoform switching, transcript boundary dynamics (including alternative promoters and polyadenylation), and transposable element-associated transcription from both PacBio and ONT datasets. By eliminating complex cross-tool data harmonization, this unified framework facilitates the transition from transcript identification to functional interpretation of transcriptomic variation. AVAILABILITY AND IMPLEMENTATION: NextLongIso is implemented in Nextflow and is freely available at github: https://github.com/YidanSunResearchLab/nf-LongIso.git and Zenodo: https://doi.org/10.5281/zenodo.21049837.

Software

Unveiling the Molecular Secrets of Seaweeds: A Comprehensive Review of Bioinformatics Applications in Algal Research.

Recent advances in high-throughput sequencing, bioinformatics, and multi-omics technologies have transformed seaweed research by overcoming long-standing challenges associated with complex genomes, diverse life cycles, and limited genomic resources. This review provides a comprehensive overview of bioinformatics approaches used to investigate seaweed genomics, transcriptomics, proteomics, metabolomics, microbiomes, and functional genomics, with emphasis on the computational tools and databases that support these analyses. Applications of bioinformatics in phylogenetics, drug discovery, microbiome characterization, and the development of biofuels, nutraceuticals, pharmaceuticals, and sustainable agriculture are also discussed. Particular attention is given to emerging strategies involving multi-omics integration, genome editing, artificial intelligence, machine learning, and synthetic biology that are reshaping seaweed research. The review further examines current challenges, including incomplete genomic resources, data standardization, and the need for experimental validation of computational predictions. Collectively, these advances highlight the growing role of bioinformatics in enabling systems-level understanding of seaweed biology and accelerating their translation into sustainable biotechnological and marine bioeconomy applications.

macroalgal genomics

Zone equalisation normalisation for improved alignment of epigenetic signal.

MOTIVATION: High-throughput genomic technologies have transformed our understanding of biological systems, yet direct comparison and visualisation of these complex datasets remains challenging. Existing normalisation methods often fail to align genomic signal across samples due to sensitivity to sequencing depth differences and localised high-signal artefacts, leading to inconsistent replicate behaviour and increased downstream variability. RESULTS: We introduce Zone Equalisation Normalisation (ZEN), a novel approach designed to improve cross-sample signal alignment of genomic data. ZEN rescales genomic signal based on variance estimated within biologically enriched regions, reducing the influence of extreme outliers while preserving underlying biological structure. Using a diverse collection of data and our new genome-wide benchmarking approach, we reveal that ZEN improves biological and technical replicate alignment across the majority of tested conditions and experimental platforms. We further show that this improved signal comparability is associated with fewer differential accessibility calls between technical replicates and a more conservative set of biological differences. Together, these results demonstrate that ZEN provides a complementary framework to improve the accuracy and reliability of genomic data analysis and that normalisation choice can affect downstream analyses and biological interpretation. AVAILABILITY AND IMPLEMENTATION: ZEN is available as an open-source Python package via conda and PyPI. Source code, documentation, tutorials, and code to reproduce the analyses are available at https://github.com/Genome-Function-Initiative-Oxford/Zone-Equalisation-Normalisation and Zenodo (https://doi.org/10.5281/zenodo.21067751).

Epigenesis, Genetic

Mitochondrial genome characteristics and phylogenetic analysis of Ramaria longispora.

This study, for the first time, assembled and annotated the complete mitochondrial genome of R. longispora using high-throughput sequencing technology. The genome is a circular molecule with a total length of 157,712 bp and a GC content of 31.55%. It encodes 71 genes, including 15 core protein-coding genes (PCGs), 25 transfer RNA (tRNA) genes, 2 ribosomal RNA (rRNA) genes, 5 free-stranding open reading frames (ORFs), and 24 intronic ORFs. Among these, most free-stranding ORFs have unknown functions but include a DNA polymerase gene, while the intronic ORFs primarily encode LAGLIDADG and GIY-YIG endonucleases. The mitochondrial genome contains 39 introns. Phylogenetic analyses based on 15 core PCGs using Bayesian inference (BI) and maximum likelihood (ML) methods revealed that this R. longispora is most closely related to Ramaria flavescens and Ramaria ichnusensis. This study provides foundational data for mitochondrial genome research in the Ramaria genus and offers important references for taxonomic and evolutionary studies of this group.

Mitochondrial genome

Expanded detection of canine enteric viruses in UK dogs with diarrhoea.

Canine enteric viruses are an important cause of gastrointestinal disease in pet dogs worldwide. Routine diagnosis often relies on pathogen-specific PCR assays, which may fail to detect some viruses, particularly neglected pathogens or genetically divergent variants of established threats. This limits both clinical characterization of affected patients and broader understanding of disease ecology. To address these limitations, we applied metagenomics and a viral discovery bioinformatics pipeline to faecal samples from diarrhoeic dogs in the UK that had been submitted routinely for PCR-based diagnostic testing. Across 80 dogs, we identified 12 viruses known to infect canids, 9 of which have not previously been reported in UK dogs. Among these, several taxa with prior associations to gastrointestinal disease were identified, including canine sapovirus and canine minute virus. By contrast, for other viruses newly detected in the UK, including bufavirus and rotavirus C, clinical relevance in dogs remains unclear. Notably, an identified protoparvovirus fell within the same species as human-canine-associated parvovirus 1, a recently described lineage detected in both canine and human oropharyngeal samples. We also identified a canine parvovirus 2 strain that clustered with a predominantly wildlife-associated lineage, consistent with occasional exposure at the domestic-wildlife interface rather than established circulation in dogs. These two detections illustrate how genome-level surveillance can help prioritize viruses for targeted investigation of host range and transmission context. Overall, these data broaden the catalogue of viruses associated with diarrhoeic dogs in the UK and support periodic review of diagnostic targets informed by viral metagenomic surveillance, while highlighting the need for controlled studies to assess causality and clinical relevance.

Animals

Genomic Profiling of Epidermal Growth Factor Receptor Mutation-Positive Non-Small Cell Lung Cancer after Progression on First-line Osimertinib: Phase II ORCHARD Study.

PURPOSE: Osimertinib is the standard of care for first-line treatment for epidermal growth factor receptor-mutated (EGFRm) non-small cell lung cancer (NSCLC). Understanding the tumor molecular profile of patients following progression on osimertinib could help inform optimal second-line treatment. PATIENTS AND METHODS: ORCHARD (NCT03944772), a phase II biomarker-directed study, enrolled patients with EGFRm NSCLC who progressed on first-line osimertinib to receive treatment based on their tumor molecular profile after progression. The study comprised three groups into which patients were allocated based on the molecular profile of their tumor, determined via next-generation sequencing (NGS) of a tumor biopsy. We report results from a prespecified, exploratory analysis of baseline tumor tissue and plasma samples to evaluate mechanisms of resistance to first-line osimertinib identified by tissue and plasma NGS. Agreement between tissue and plasma NGS data was also assessed. RESULTS: This study provided a comprehensive dataset exploring tissue (n = 400) and plasma (n = 191) genomics, enabling characterization of the histogenomic landscape after first-line osimertinib treatment. TP53 and MDM2/4 alterations were mutually exclusive and occurred in 86% of tumors. When combining tissue and plasma genomics, resistance alterations were detected in 87% of samples, with multiple resistance alterations in 46%. Alterations in the PI3K pathway, SOX2, and MYC were frequently detected in histologically transformed tumors. Additionally, differential patterns of co-occurring EGFR mutations in tumors with L858R versus exon 19 deletion were observed. CONCLUSIONS: This comprehensive analysis highlights potential heterogeneous resistance to first-line osimertinib treatment, providing a rationale for combining treatments with broad activity to improve patient outcomes. See related commentary by Gupta et al., p. 3718.

Humans

A novel relationship between time offsets in capillary electrophoresis and DNA sequence variations in short tandem repeats.

Next-generation sequencing (NGS) provides increased discriminatory power in forensic DNA analysis due to the detection of isoalleles. Differences in sequences between alleles allow for a second layer of differentiation between DNA contributors beyond the number of short tandem repeat (STR) repeat units. However, because NGS is a more time and resource-intensive analysis than conventional capillary electrophoresis (CE), laboratories may benefit from indicators that suggest NGS is likely to provide added value. This study examined whether CE migration offsets, measured as residuals in the OSIRIS analysis software, can differ significantly among STR isoalleles. Residuals represent the time offset between a sample allele peak and its corresponding allelic ladder peak. Paired CE and NGS data from 95 single source samples were analyzed for CE-based residual differences, as the NGS data provided the sequence information of the corresponding isoalleles. Residual values differed significantly among isoalleles at several STR loci. Statistically significant differences were identified at D16S539 and D3S1358, as well as at specific allele lengths within D12S391, D13S317, and D8S1179. These findings demonstrate that CE residual variation can reflect underlying STR sequence differences between contributors. In practice, residual-based metrics could help laboratories to identify casework reference samples where NGS is likely to provide additional discrimination, without the need for processing outside of a routine CE workflow. Due to the potentially large number of isoalleles, community wide efforts to aggregate CE residual differences versus isoallele sequences may be useful in the validation and implementation of this approach to add value to forensic DNA analyses.

Electrophoresis, Capillary

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software

Comprehensive Viral Detection and Profiling of Plasma Cell-Free RNA in Patients With Suspected Hemophagocytic Lymphohistiocytosis.

Hemophagocytic lymphohistiocytosis (HLH) is a severe, rapidly progressive disease. While viral infection is considered a common etiology of pediatric HLH, specific causative viruses other than the Epstein-Barr virus (EBV) have been rarely identified. This study utilized metagenomic next-generation sequencing (NGS) to identify potential causative pathogens in plasma samples from 17 pediatric patients with suspected HLH. Additionally, one case each of confirmed EBV- and cytomegalovirus (CMV)-associated HLH was analyzed for methodological validation. Plasma cell-free RNA (cfRNA) profiling was performed using NGS data to assess the host transcriptome response. Significant viral reads of human herpesvirus-6B, human herpesvirus-7, and Hubei reo-like virus (HRLV) 14 were detected using metagenomic NGS in one patient each. Plasma cfRNA profiles from five patients with viral infection (including EBV and CMV) were compared to those of 14 patients without viral infection. By comparing the two patient groups, 1053 differentially expressed genes were identified. The gene ontology (GO) term of "adaptive immune response" (GO: 0002250) was significantly enriched among upregulated genes in the virus-positive group. Furthermore, an isolated cluster consisting specifically of mitochondrial RNAs, was identified in the upregulated genes of the virus-positive group. Using metagenomic NGS, several candidate viral pathogens were identified in patients with suspected infection-related HLH. The viral genome of HRLV 14, previously undetected in human clinical samples, was identified in one patient. The results from plasma cfRNA profiling suggest that mitochondrial RNAs may reflect the underlying pathogenesis of virus-associated HLH and have potential utility as disease biomarkers.

Humans

Accurate somatic small variant discovery for multiple sequencing technologies with DeepSomatic.

Somatic variant detection is an integral part of cancer genomics analysis. While most methods have focused on short-read sequencing, long-read technologies offer potential advantages in repeat mapping and variant phasing. We present DeepSomatic, a deep-learning method for detecting somatic small nucleotide variations and insertions and deletions from both short-read and long-read data. The method has modes for whole-genome and whole-exome sequencing and can run on tumor-normal, tumor-only and formalin-fixed paraffin-embedded samples. To train DeepSomatic and help address the dearth of publicly available training and benchmarking data for somatic variant detection, we generated and make openly available the Cancer Standards Long-read Evaluation (CASTLE) dataset of six matched tumor-normal cell line pairs whole-genome sequenced with Illumina, PacBio HiFi and Oxford Nanopore Technologies, along with benchmark variant sets. Across samples, both cell line and patient-derived, and across short-read and long-read sequencing technologies, DeepSomatic consistently outperforms existing callers.

Humans

Targeted sequencing reveals a distinct genetic alteration landscape in oral multiple primary squamous cell carcinomas.

OBJECTIVE: Oral multiple primary cancers (MPCs) are associated with poor clinical outcomes, yet their genomic characteristics remain insufficiently understood. DESIGN: Fifty-four formalin-fixed paraffin-embedded (FFPE) tumor samples from 30 patients with oral MPCs were analyzed using high-depth targeted sequencing of a customized 14-gene panel derived from prior whole-exome sequencing data. Detected alterations were analyzed after removal of synonymous mutations. RESULTS: Non-silent genomic alterations were identified in 59.3% (32/54) of samples, involving 19 patients. A total of 70 variant loci across 13 genes were detected. AKAP13 was the most frequently mutated gene at both the sample (22.2%, 12/54), with recurrent mutations observed across multiple patients. In contrast, TP53 mutations occurred at a substantially lower frequency (11.1%, 6/54). Marked inter- and intra-patient mutational heterogeneity was observed. CONCLUSIONS: FFPE-based targeted sequencing enabled an initial characterization of genomic alterations in oral MPCs. Recurrent alterations in AKAP13, GLI2, JMJD1C, and DNAH8, together with the relatively low frequency of TP53 alterations, identify candidate genomic features for further investigation and provide a basis for future studies of the molecular basis of oral MPCs.

Humans

Upscaling Genotyping by Amplicon Sequencing With GBAS-GUI.

Genotyping by amplicon sequencing (GBAS) is a relatively low-cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large-scale genetic monitoring projects. However, most existing analytical pipelines are either marker-specific, insufficiently scalable, or lacking efficient data management systems for the long-term integration of genotypic information, limiting the full potential of GBAS. Here, we address this gap by introducing GBAS-GUI (https://github.com/sonnenbe-dot/GBAS-GUI), a pipeline capable of generating GBAS-based genotypic data for a wide variety of loci at scale. GBAS-GUI integrates a graphical user interface with multiple checkpoints to improve accessibility and robustness. It implements multiprocessing architecture and a relational database that links genotypic data with associated sample metadata to enhance scalability and data management. The pipeline further enables marker screening through automated calculation of polymorphism information content (PIC) and implements a strategy to recover homologous genotypic information from paralogous loci with non-overlapping amplicon length ranges. Using multiple empirical datasets, we demonstrate substantial improvements in processing speed, database management and handling artefacts related to co-amplification of unspecific regions and duplicates of the same genomic region. We further show that incorporating the full sequence information captured by an amplicon increases marker information content beyond what is achievable with length-based genotyping alone and expands the analytical versatility of GBAS. Overall, GBAS-GUI provides a robust, scalable and versatile framework that unlocks the potential of GBAS for large-scale population genetic and phylogeographic studies.

Genotyping Techniques

Sequencing approaches in hereditary cancer testing: strengths, limitations and future directions.

Over the past three decades, Hereditary Cancer Testing (HCT) has evolved from single gene assays into multigene panel testing (MGPT), which allows for the screening of all known hereditary cancer genes in a single assay. MGPT is currently the standard approach for clinical HCT. However, with decreasing sequencing costs and increased instrument throughput, the scalability of exome sequencing (ES) and genome sequencing (GS) for HCT indications is becoming more viable. These methods provide broader insights into the coding exons and/or the entire genome, respectively. ES/GS data can also be reanalyzed to identify variants in novel genes that were not characterized at the time of initial testing, or to support research efforts aimed at uncovering additional associations between germline variants and cancer predisposition. Additionally, the emerging use of long-read sequencing (LRS) is noteworthy, enabling improved variant detection compared to short-read sequencing, especially for complex/structural variants and variation in difficult-to-sequence or paralogous regions in genes such as PMS2. This has the potential to increase the accuracy of HCT, reduce the turnaround time, find previously unidentifiable cancer risk variants, and ultimately increase the diagnostic yield. This article provides a comprehensive summary of the sequencing approaches used in HCT, discussing their strengths and limitations. We also highlight the added value of complementing DNA-only testing with RNA and tumor sequencing. Furthermore, we explore LRS-based approaches and discuss opportunities for their implementation in routine genetic testing for hereditary cancer.

Humans

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans

Optimizing GRIDSS for clinical use: A targeted NGS filtering strategy for germline structural variant detection.

Detecting intermediate-sized structural variants (SVs) remains challenging in diagnostics, as tools for single-nucleotide and copy-number variants, particularly read-depth-based methods, are often insufficient. GRIDSS addresses this gap by integrating paired-end mapping, split-read analysis, and assembly-based approaches. However, its use in targeted sequencing and diagnostic workflows remains complex. NGS panel data from 9726 patients with suspected hereditary cancer were analyzed using GRIDSS. A filtering strategy was developed to prioritize clinically relevant germline SVs. Multiple parameter settings were tested to optimize performance. The initial dataset of 1,307,592 variants was reduced to 89 candidates after applying the selected filtering strategy. Of these, 24 had been previously detected by routine callers and were not further analyzed. Among the remaining 65, 13 were considered likely true positives after visual inspection using IGV. Experimental validation was performed by Sanger/Nanopore long-read sequencing for these variants, all of which were confirmed. Eight were classified as (likely) pathogenic, including two frameshift duplications in MSH6, one splicing variant in BARD1, and five mobile element insertions in APC, BRCA2, and PALB2. Altogether, GRIDSS implementation increased diagnostic yield while maintaining feasibility for diagnostic workflows. Comprehensive workflow scheme for germline structural variant detection and results in our diagnostic setting.

Humans

Lynch syndrome-associated urothelial carcinoma: clinical and molecular findings from a single-institution cohort.

Lynch syndrome-associated urothelial carcinoma (LS-UC) is a rare and undercharacterized clinical entity. While FGFR3 alterations are well described in sporadic urothelial carcinoma, their prevalence and clinical implications in LS-UC remain unclear. We aimed to provide a comprehensive clinical and molecular characterization of LS-UC. We conducted a retrospective single-center study including patients with Lynch syndrome (LS) and histologically confirmed urothelial carcinoma (UC). Clinical, pathological, treatment, and follow-up data were collected. Targeted next-generation sequencing was performed on available tumor samples to assess genomic alterations, with particular attention to FGFR3 mutations. A total of 27 patients with LS-UC were identified, with a predominance of upper urinary tract involvement (70%). Most tumors were diagnosed at an early stage and initially managed with local treatment. During a median follow-up of 92 months, 48% of patients experienced recurrence, with a median time to recurrence of 37 months. Recurrences were predominantly local and were mainly managed with additional surgical or intravesical treatments. No deaths were attributable to UC at last follow-up. Molecular analysis was feasible in 9 cases. FGFR3 mutations were detected in 67% of evaluable samples, with the recurrent p.Arg248Cys hotspot identified in 55% of cases. Additional alterations involved TP53, SWI/SNF complex genes, and PIK3CA, which co-occurred with FGFR3 p.Arg248Cys. No gene fusions were identified. This study expands the limited molecular and clinical evidence on Lynch syndrome-associated urothelial carcinoma. Beyond confirming the recurrent role of FGFR3 (notably p.Arg248Cys), our comprehensive multigene profiling enriches the current genomic knowledge for this rare population. Multi-center collaborative efforts remain essential to aggregate larger datasets and ultimately guide personalized patient management.

Humans

Obtaining a Diagnostic Yield via Scan findings prior to the introduction of SEquencing retrospectivelY (ODYSSEY): a cohort study.

OBJECTIVE: To determine the retrospective yield of prenatal exome sequencing (PES) by establishing the proportion of children with a postnatal monogenic diagnosis that could have been diagnosed prenatally if PES had been available. METHODS: The study cohort comprised a sample of children in Northern Ireland, born between January 2010 and January 2018 (predating routine availability of PES), who received a monogenic diagnosis postnatally via next generation sequencing as part of either of two UK-wide studies (the 100 000 Genomes Project (2015-2018) or the Deciphering Developmental Disorders study (2011-2015)). Clinical data were collected retrospectively and correlated with the current UK National Health Service PES protocol, including the phenotypic eligibility criteria for PES and the associated fetal anomalies gene panel. Cases were considered retrospective diagnoses if the fetal phenotype would have been eligible for PES and the diagnostic gene was included on the test panel, meaning prenatal diagnosis in this current era could have been feasible. RESULTS: Of 101 children, 17.8% (95% CI, 10.3-25.3%) had both an eligible fetal structural anomaly (FSA) (i.e. high-risk FSA) and a diagnostic gene on the associated test panel, meaning that they could have been diagnosed prenatally in the current clinical landscape. The median length of the diagnostic odyssey for this subgroup of children was 3.7 years (1354 (range, 822-2450) days). Moreover, 58.4% (n = 59) of cases had no anomalies detected prenatally and 19.8% (n = 20) had a FSA that would not meet the eligibility criteria for PES (low-risk FSA). Although these cases would have been ineligible for PES under the current clinical pathway, 89.9% (n = 71/79) were affected by severe or profound syndromes. Postnatally, the most common functional anomalies were neurodevelopmental delay/intellectual disability and/or behavioral abnormality, which were observed in 80.2% (n = 81) of the included children. However, 80.2% (n = 65/81) of these affected children did not present with fetal anomalies eligible for PES. CONCLUSIONS: Almost one-fifth of children with a monogenic condition included in this study could have received a diagnosis via modern PES, avoiding a diagnostic odyssey lasting almost 4 years. However, despite having a monogenic condition, over half of the children did not present with any structural anomalies in utero. This demonstrates the degree to which fetal imaging is limited in its ability to reassure parents of the absence of a fetal genetic syndrome. © 2026 The Author(s). Ultrasound in Obstetrics & Gynecology published by John Wiley & Sons Ltd on behalf of International Society of Ultrasound in Obstetrics and Gynecology.

Humans

Comparative evaluation of probe-capture and conventional metagenomic sequencing across multiple clinical sample types, with analysis of paired bronchoalveolar lavage fluid and blood samples.

Conventional metagenomic next-generation sequencing (mNGS) suffers from host nucleic acid interference and poor performance in low-biomass samples. Probe-capture metagenomic sequencing (PC-mNGS), which enriches microbial targets via hybridization probes, shows superior sensitivity but lacks systematic multi-sample evaluations. This study compared PC-mNGS and mNGS across diverse clinical specimens (bronchoalveolar lavage fluid [BALF], blood, cerebrospinal fluid [CSF]) and assessed the clinical utility of pathogen co-detection in paired BALF-blood samples from sepsis patients. A total of 282 samples (81 BALF, 141 blood, 25 CSF, 35 others) sequenced by both PC-mNGS and mNGS were analyzed. Additionally, 621 paired BALF-blood samples from sepsis patients with pulmonary infections were evaluated. PC-mNGS achieved higher pathogen detection rates (66.67% vs 57.10%, P = 0.000198) than mNGS, particularly in blood (66.67% vs 47.52%, P = 2.5 × 10⁻⁵). PC-mNGS detected more bacteria (19 species exclusive) and fungi (11 species exclusive) than mNGS. Viruses showed comparable detection. BALF and CSF exhibited high overall agreement (OPA: 96.30% and 88%, respectively), while blood had lower concordance (NPA: 54.05%, OPA: 70.92%). A total of 60.55% of BALF-positive samples (PC-mNGS) had co-detected pathogens in blood. Gram-negative bacteria (e.g., Klebsiella pneumoniae) and fungi (e.g., Candida albicans) showed higher blood co-detection rates than viruses. In this study, PC-mNGS detected more pathogens and showed a higher positivity rate than mNGS in blood samples. BALF sequencing data, particularly bacterial reads per million (RPM), may predict bloodstream co-detection, aiding in sepsis management. However, clinical validation and integration with traditional diagnostics are needed to confirm utility. This study highlights PC-mNGS as a promising tool for complex infections but underscores the need for rigorous multi-context validation.IMPORTANCEAccurate and rapid identification of pathogens is critical for effective treatment of severe infectious diseases, such as sepsis. This study demonstrates that probe-capture metagenomic sequencing (PC-mNGS) detected more pathogens in blood samples compared to conventional metagenomic sequencing, especially for bacterial and fungal infections. By analyzing paired lung and blood samples, we show that high pathogen levels in lung fluid may predict bloodstream infection, offering a potential early warning for clinicians. These findings support the use of PC-mNGS as a more sensitive diagnostic tool, which could lead to faster, more targeted therapies and better outcomes for patients with complex infections.

Humans