PubMed HealthSearch

SEARCH · PubMed Health

Results for “Whole genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Transmission of extended spectrum β-lactamase-producing Escherichia coli and antimicrobial resistance gene flow across One Health compartments in eastern Africa: a whole-genome sequence analysis from a prospective cohort study.

BACKGROUND: The One Health paradigm considers interdependence of human, animal, and environmental health. However, there is little evidence from high-income countries to support the importance of a One Health approach to addressing spread of antimicrobial resistance (AMR). Given AMR is a global threat, understanding how the close interactions of humans with animals and the environment in low-income settings affect the spread of AMR is important. We aimed to investigate diversity and transmission of extended spectrum β-lactamase (ESBL)-producing Escherichia coli across household-linked One Health compartments using genomic data. METHODS: We sequenced whole genomes of ESBL-producing E coli isolates from humans, animals, and the environment from a prospective, longitudinal cohort study conducted in Malawi (April 29, 2019, to Dec 3, 2020) and Uganda (July 16, 2020, to Aug 6, 2021). In the cohort study, 259 households were enrolled at baseline in Malawi and 92 in Uganda from a mix of urban, peri-urban, and rural areas. Households were followed up at months 1, 3, and 6 in Malawi and at months 1, 2, and 4 in Uganda. Samples collected at each visit included human and animal stool, environmental samples from hand-contact areas, food, and water, and broader environmental samples such as river water. Samples were cultured in buffered peptone water and then ESBL chromogenic agar to isolate ESBL-producing E coli. ESBL-producing E coli isolates underwent whole-genome sequencing. We performed phylogenetic analyses, and in-silico multi-locus sequence typing, characterised AMR determinants and linked genotypes with sample location, ecological source, and other covariates. We performed fine-scale single nucleotide polymorphism (SNP) and network analysis to infer strain and plasmid transmission across ecological compartments. The primary outcome was colonisation with ESBL-producing E coli. Secondary outcomes were genomic clusters and ESBL genomic determinants within and between One Health compartments. FINDINGS: We found high diversity of ESBL-producing E coli, with 170 sequence types and 166 genomic clusters identified from 2344 genomes, including 1814 genomes from Malawi (907 human, 221 animal, and 686 environmental) and 530 genomes from Uganda (380 human, 147 animal, and three environmental). Sequence type (ST)131 dominated in Malawi (209 [11·5%] of 1814 genomes), and ST10 dominated in Uganda (45 [8·5%] of 530 genomes). Common ESBL genes blaCTX-M-15 (1604 [68·4%] of 2344 genomes) and blaCTX-M-27 (336 [14·3%] of 2344 genomes) were carried on a complex network of 55 and 30 different plasmids. This diversity of plasmids presented multiple pathways for dissemination and revealed high force of selection. Phylogenetic analyses revealed common intermixing of isolates between humans, animals, and the environment. SNP transmission analysis revealed ecologically overlapping clusters, suggesting ESBL-producing E coli co-circulation both within and between compartments with frequent spillover events. Applying a five-SNP threshold, we inferred 463 human-environment transmission events, 146 human-animal events, and 142 animal-environment events. INTERPRETATION: Our work suggests that a One Health approach is crucial to addressing AMR in eastern Africa. Improving water, sanitation, and hygiene systems will create a safer environment, reduce spillovers of AMR bacteria between compartments, and eventually reduce AMR reservoirs in the environment and in animals. FUNDING: Medical Research Council, National Institute for Health and Care Research, and Wellcome Trust.

Humans

Whole-Genome Sequencing Reveals Population Structure, Genetic Diversity, and Selection Signatures in Kazakh Dromedary and Bactrian Camels.

Understanding the genomic basis of environmental adaptation is essential for the conservation and genetic improvement of domestic camels. In this study, we investigated the population structure, genetic diversity, and genomic variation potentially associated with environmental adaptation of Kazakh dromedary and Bactrian camels using whole-genome sequencing. Whole-genome sequencing data were generated for Kazakh camels (15 dromedaries and 16 Bactrian camels) and integrated with 131 publicly available genomes representing camel populations from the Arabian Peninsula, Iran, Xinjiang, Inner Mongolia, and Mongolian wild camels. Population structure, genetic diversity, and genome-wide selection were evaluated using principal component analysis, ADMIXTURE, nucleotide diversity, linkage disequilibrium, runs of homozygosity, genomic inbreeding (FROH), and selection scans based on FST, θπ ratio, and XP-EHH. Population genomic analyses revealed clear differentiation between dromedary and Bactrian camels, whereas Kazakh camel populations exhibited higher nucleotide diversity (θπ = 1.307-1.551 × 10-3), and lower genomic inbreeding (median FROH: 0.037-0.056) than Arabian populations. Genome-wide selection analyses identified MC4R as the prominent candidate gene in Kazakh dromedaries and RYR1 as a prominent candidate gene in Kazakh Bactrian camels. Functional enrichment analyses highlighted pathways related to energy metabolism, thermogenesis, calcium signaling, skeletal muscle function, mitochondrial activity, and oxidative stress response. These findings provide new insights into genomic variation potentially associated with environmental adaptation in Kazakh camels and offer valuable genomic resources for future conservation, breeding, and evolutionary studies.

MC4R

ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing.

MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1-5×, sWGS data display a pronounced zero‑inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.

DNA Copy Number Variations

Molecular residual disease assessment in colorectal and bladder cancer by somatic structural variant analysis of cell-free DNA whole-genome sequencing data.

BACKGROUND: Whole-genome sequencing (WGS)-based methods for circulating tumor DNA (ctDNA) detection typically rely on tumor-informed identification of somatic single nucleotide variants (SNVs). Somatic structural variants (SVs) are another type of cancer-specific genomic alteration, which owing to their larger genomic footprint and unique breakpoint junctions, are easier to distinguish from sequencing noise than SNVs. They are, however, rarely used for ctDNA detection because of (1) artifacts from WGS procedures that SV callers may falsely interpret as genuine SVs. This makes it difficult to establish high-confidence SV catalogos from short-read tumor WGS and can cause false-positive ctDNA detections. (2) Lack of robust strategies to quantify SV-supporting reads in plasma WGS. To address these barriers and enable integration of SV biomarkers into WGS-based ctDNA detection, we present a bioinformatic framework for algorithmic curation of somatic SV calls from fresh-frozen and formalin-fixed paraffin-embedded (FFPE) tumors, coupled with a novel approach for sensitive, accurate mapping and quantification of SV breakpoint-supporting reads in plasma WGS. METHODS: Tumor, normal and plasma WGS data from 144 patients with stage III colorectal cancer was used to establish the bioinformatic framework. This included ~30x WGS data from 1564 serially collected plasma samples. The framework was validated using tumor/normal/plasma WGS data from 32 patients with muscle-invasive bladder cancer. SV-based ctDNA detection was benchmarked against previously published SNV-based ctDNA results for the same samples. RESULTS: After curation of SV calls and quantification in plasma WGS, our SV-based approach enabled robust ctDNA detection with overall specificity exceeding 99% in plasma samples. Furthermore, we observed strong concordance (Pearson&#x2019;s r&#x2009;>&#x2009;0.93, p&#x2009;<&#x2009;2.2&#x2009;&#xd7;&#x2009;10&#x2212; 16) between ctDNA-positive samples identified by our SV-based method and previous SNV-based analyses, validating the reliability of our approach. Finally, we demonstrated application of the method in an independent bladder cancer cohort, highlighting its generalizability and potential clinical use. CONCLUSIONS: We provide a bioinformatic framework that establishes somatic SVs as ultra-specific biomarkers for WGS-based, tumor-informed ctDNA detection. The approach delivers specific detection even when the SV catalogos are established from FFPE samples. The SV framework can stand alone or enhance SNV-based analysis pipelines.

Humans

META-DIFF: a k-mer-based pipeline that detects differentially abundant sequences in metagenomics whole genome sequencing.

Traditional case-control metagenomic studies are constrained by their dependence on taxonomic and functional databases. Because annotation occurs before differential analysis, they are limited to known elements and keep function and taxonomy separate. Although binning strategies have emerged to reconstruct genomes and mitigate this issue, they still require an assembly step, preventing the use of all available sequencing data. Here, we introduce META-DIFF, a pipeline based on differentially abundant k-mers independently of any prior annotation. From those k-mers, it reconstructs longer sequences and provides biological context, as well as the best set of unitigs to discriminate between conditions. Across both taxonomy-centric and functionally-centric benchmarks, it showed robust performance and displayed great reproducibility. It also behaved more conservatively than did other univariate methodologies, i.e. it maintained a high precision at the expense of recall, particularly in conditions of low fold-change and limited sequencing depth. The efficacy of META-DIFF was further validated through its application to a real-world colorectal cancer dataset, which produced both confirmatory and novel results compared with those of previous publications. The pipeline is able to exploit all reads and identify differentially abundant elements, including unknown DNA, prior to annotation. With the guidelines provided, META-DIFF provides users with great exploratory power to unravel microbiome changes.

Metagenomics

Whole-genome sequencing of 490,640 UK Biobank participants.

Whole-genome sequencing provides an unbiased and complete view of the human genome and enables the discovery of genetic variation without the technical limitations of other genotyping technologies. Here we report on whole-genome sequencing of 490,640 UK Biobank participants, building on previous genotyping effort1. This advance deepens our understanding of how genetics associates with disease biology and further enhances the value of this open resource for the study of human biology and health. Coupling this dataset with rich phenotypic data, we surveyed within- and cross-ancestry genomic associations and identified novel genetic and clinical insights. Although most associations with disease traits were primarily observed in individuals of European ancestries, strong or novel signals were also identified in individuals of African and Asian ancestries. With the improved ability to accurately genotype structural variants and exonic variation in both coding and UTR sequences, we strengthened and revealed novel insights relative to whole-exome sequencing2,3 analyses. This dataset, representing a large collection of whole-genome sequencing&#xa0;data that is available to the UK Biobank research community, will enable advances of our understanding of the human genome, facilitate the discovery of diagnostics and&#xa0;therapeutics with higher efficacy and improved safety profile, and enable precision medicine strategies with the potential to improve global health.

Humans

A tiled amplicon protocol for culture-free whole-genome sequencing of M. tuberculosis from clinical specimens.

Whole-genome sequencing of Mycobacterium tuberculosis can be a valuable tool for TB surveillance and treatment, providing insights into transmission patterns and comprehensive drug susceptibility testing. However, the slow growth of M. tuberculosis means traditional culture-based sequencing methods can take weeks to return results, which has limited the widespread adoption of these techniques and limited their use in clinical decision-making. Tiled amplicon sequencing is a fast, reliable, and cost-effective method of whole-genome sequencing that can be done directly on clinical specimens and has been implemented at scale in academic and public health laboratories across the world; it was the cornerstone of SARS-CoV-2 sequencing and has been adapted for a wide range of viral pathogens. However, similar methods are not yet available for far larger bacterial genomes. Extending this approach to M. tuberculosis would significantly reduce the cost, labor, and turnaround time for whole-genome sequencing. We designed a tiled amplicon panel consisting of 5,128 primers that covers the entire M. tuberculosis genome, the largest tiled amplicon sequencing panel we are aware of to date. Applying our amplicon panels to clinical samples of sputum, we show the ability to recover whole-genome bacterial sequences without the need for culture. The resulting sequence data can be used to determine M. tuberculosis lineage and reliably identify markers of drug resistance. Using this approach in clinical settings could reduce the time needed for comprehensive drug susceptibility testing from weeks to days and enable genomic epidemiology to be performed at scale, even in resource-limited settings.IMPORTANCEWe have developed and tested an amplicon panel, TB-seq, for the priority pathogen Mycobacterium tuberculosis, demonstrating recovery of near-full genomes directly from patient sputum, including mixed and low-concentration samples. This approach significantly reduces the turnaround time for this slow-growing bacterium while maintaining high accuracy in detecting clinically relevant mutations, including those associated with drug resistance. Given the global burden of tuberculosis and the critical need for faster diagnostic solutions, we believe our method has the potential to improve clinical decision-making and public health strategies.

Mycobacterium tuberculosis

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans

Whole genome sequencing and phylogenetic classification accelerate the implementation of respiratory syncytial virus genomic surveillance in Canada: a pilot study.

UNLABELLED: Whole genome sequencing (WGS) has emerged as a powerful tool to facilitate the study of existing and emerging infectious diseases. WGS-based genomic surveillance provides information on the genetic diversity and tracks the evolution of important viral pathogens, including respiratory syncytial virus (RSV). Multiplex tiling polymerase chain reaction (PCR) assays have been used to facilitate sequencing of a variety of pathogens in support of genomics-based surveillance initiatives. We developed, optimized, and implemented multiplex tiling PCR assays for RSVA and RSVB capable of generating near-complete genomes in the majority of contemporaneous specimens tested. A pilot data set comprising 52 RSVA and 37 RSVB genomes derived from Canadian clinical specimens during the 2022-2023 respiratory virus season was used to perform phylogenetic analyses using both near-complete genome and glycoprotein (G) sequences. Overall, the RSV phylogenetic tree built with whole genomes showed identical lineage clusters as compared to the G gene but was more discriminatory. Moreover, the availability of complete genomes enables the identification of a broader range of mutations. For instance, mutations identified in the fusion protein among Canadian isolates tested here, including S377N, K272M, S276N, S211N, S206I, and S209Q, could affect the efficacy of current vaccines or antiviral-based therapeutics. In conclusion, our work reinforces other recent studies demonstrating the utility of multiplex tiling PCR assays to facilitate high-throughput WGS of RSV, which is capable of supporting enhanced genomic surveillance initiatives, as well as the more comprehensive genomic analyses required to inform public health strategies for the development and usage of vaccines and antiviral drugs. IMPORTANCE: We present assays to efficiently sequence genomes of RSVA and RSVB. This enables researchers and public health agencies to acquire high-quality genomic data using rapid and cost-effective approaches. Genomic data-based comparative analysis can be used to conduct surveillance and monitor circulating isolates for efficacy of vaccines and antiviral therapeutics.

Humans

Routine methods misidentify Serratia spp.: Limitations of MALDI-TOF MS revealed by whole-genome sequencing.

Accurate species-level identification within the genus Serratia remains challenging due to extensive phenotypic overlap and high genomic relatedness among closely related and recently described taxa. This study presents an evaluation of routine and genome-based identification approaches applied to clinical Serratia isolates, integrating phenotypic assays, MALDI-TOF MS (Bruker Daltonics), 16S rRNA gene sequencing, and Whole-Genome Sequencing (WGS). A total of 103 isolates collected from a teaching hospital were analyzed. WGS was performed on a subset of isolates. Conventional biochemical methods classified all isolates as Serratia marcescens, whereas MALDI-TOF MS identified 60.1% as S. marcescens, 11.6% as S. ureilytica, and 28.1% just at the genus level. Peak analysis from MALDI-TOF MS revealed specific peaks associated with S. marcescens and S. ureilytica, but limited discriminatory power. WGS of six isolates initially identified as S. ureilytica by MALDI-TOF MS revealed reclassification as Serratia sarumanii (n = 5) and Serratia montpellierensis (n = 1), supported by Average Nucleotide Identity (ANI), Average Amino Acid Identity (AAI), and Digital DNA-DNA Hybridization (dDDH) thresholds. In contrast, 16S rRNA analysis showed limited species-level resolution. Phylogenomic and SNP-based analyses confirmed these classifications with strong support. Overall, this study underscores the critical role of high-resolution genomic approaches for precise species identification and highlights the need for continuous expansion and curation of MALDI-TOF MS reference databases to support reliable clinical diagnostics and epidemiological surveillance of emerging Serratia species.

Spectrometry, Mass, Matrix-Assisted Laser Desorpti

Large-scale simulation of coverage and error rate tradeoffs for cancer detection in cell-free DNA whole-genome sequencing.

MOTIVATION: Cell-free DNA (cfDNA) whole-genome sequencing (WGS) is a promising approach for detecting cancer recurrence. It enables cancer detection by identifying all tumor-derived cfDNA (ctDNA) molecules carrying somatic single nucleotide variants (sSNVs). While ideally, a sequencing platform should be highly accurate for reliable ctDNA detection, in reality, all sequencing platforms introduce sequencing errors that generate false positives indistinguishable from true SNVs. Understanding how sequencing parameters influence ctDNA detection sensitivity at low tumor fractions (TFs) in cfDNA samples is essential for guiding sequencing strategies in clinical contexts. To model cfDNA sequencing for tumor detection, which contains asymmetric noise and multiple interacting parameters, analytical modeling is intractable, motivating large-scale parallelized simulation. RESULTS: We developed a simulation framework to generate in silico cfDNA data across 10 cancer types. In total, 480 million cfDNA samples were simulated from tumor WGS profiles. Overall, the lowest detectable TF differs substantially between cancer types under identical sequencing conditions due to variations in mutational load. For cancers with high mutational load, 3&#xd7; coverage with low-error techniques reliably detects TFs below 0.1%. In contrast, cancers with low mutational load require at least six-fold higher coverage to achieve comparable detection thresholds. Increasing sequencing quality scores from Q30 to Q55 at 30&#xd7; coverage further enhances sensitivity, enabling detection of TFs as low as 1&#x2009;&#xd7;&#x2009;10-5. This study provides a comprehensive framework for optimizing sequencing parameters, offering valuable guidance for tailoring future technology development for specific cancer types and clinical applications. AVAILABILITY AND IMPLEMENTATION: The code is publicly available at https://github.com/UMCUGenetics/cfdetect/tree/main.

Whole Genome Sequencing

Genetic structure and demographic history of house mice in western Europe inferred using whole-genome sequences.

The western house mouse, Mus musculus domesticus, is a human commensal and an outstanding model organism for studying a wide variety of traits and diseases. However, we have few genomic resources for wild mice and only a rudimentary understanding of the demographic history of house mice in Europe. Here, we sequenced 59 whole genomes of mice collected from England, Scotland, Wales, Guernsey, northern France, Italy, Portugal and Spain. We combined this dataset with 24 previously published sequences from southern France, Germany and Iran and compared patterns of population structure and inferred demographic parameters for house mice in western Europe to patterns seen in humans. Principal component and phylogenetic analyses identified three genetic clusters in western European mice. Admixture and f-branch statistics identified historical gene flow between these genetic clusters. Demographic analyses suggest a shared history of population bottlenecks prior to 20&#x202f;000 years ago. Estimated divergence times between populations of house mice from western Europe ranged from 1500 to 5500 years ago, in general agreement with the zooarchaeological record. These results correspond well with key aspects of contemporary human population structure and the history of migration in western Europe, highlighting the commensal relationship of this important genetic model.

Animals

Whole genome sequence of a superbug-Escherichia coli strain KAB-AI-497 isolated from the vagina of a 20 year old pregnant woman with premature rupture of membrane (PROM) in a resource limited setting, Kabale Regional Referral Hospital, in Uganda.

OBJECTIVES: The objective of the study is to sequence the whole genome of multidrug resistant E. coli strain KAB-AI-497 that causes bacterial vaginosis and implicated in premature rupture of membrane in pregnant woman. DATA DESCRIPTION: The DNA of the E. coli strain KAB-AI-497 was extracted using the MagAttract HMW DNA Kit, and the extracted DNA was sequenced using an MGI DNBSEQ G99ARS platform. FastQC was used to perform quality control analysis and the reads were trimmed by Trimmomatic. De novo genome assembly was performed by SPAdes and it resulted to a draft assembled genome that has 5.1 Mb genome size, 153 contigs, and 50.5% GC content. Quality analysis of the assembled genome revealed it has 98.46% completeness and 0.97% contamination. The closest E. coli strain to this strain KAB-AI-497 in terms of similarity was Escherichia coli SMS-3-5 with an average nucleotide identity of 98.43% and genome coverage of 86.19%, which confirmed the species level identity of the strain. The assembled genome was annotated using the NCBI Prokaryotic Genome Annotation Pipeline which identified 4,726 protein coding genes in the strain genome. Furthermore, the annotation revealed the genome has resistant genes responsible for resistance against many antibiotic classes such as tetracycline, fluoroquinolone, and penicillin.

Escherichia coli

Whole genome sequence analysis of low-density lipoprotein cholesterol across 246&#xa0;K individuals.

BACKGROUND: Rare genetic variation provided by whole genome sequence datasets has been relatively less explored for its contributions to human traits. Meta-analysis of sequencing data offers advantages by integrating larger sample sizes from diverse cohorts, thereby increasing the likelihood of discovering novel insights into complex traits. Furthermore, emerging methods in genome-wide rare variant association testing further improve power and interpretability. RESULTS: Here, we conduct the largest meta-analysis of whole genome sequencing for low-density lipoprotein cholesterol (LDL-C), a therapeutic target for coronary artery disease, analyzing data from 246&#xa0;K participants and integrating 1.23B variants from the UK Biobank and the Trans-Omics for Precision Medicine (TOPMed) program. We identify numerous rare coding and non-coding gene associations related to LDL-C, with replication across 86&#xa0;K participants in All of Us. Our findings are based on single-variant analyses, rare coding and non-coding variant aggregation tests, and sliding window approaches. Through this comprehensive analysis, we identify 704 novel single-variant associations, 25 novel rare coding variant aggregates, 28 novel rare non-coding variant aggregates, and one novel sliding window aggregate. CONCLUSIONS: This study provides a meta-analysis framework for large-scale whole genome sequence association analyses from diverse population groups, yielding novel rare non-coding variant associations.

Humans

RUMINA: high-throughput deduplication of unique molecular identifiers for amplicon and whole-genome sequencing with enhanced error correction.

MOTIVATION: Unique molecular identifiers (UMIs) are widely used in next-generation sequencing to enable accurate molecular counting and error correction. However, challenges remain in accurately collapsing UMI clusters, especially when read counts are low or sparse read clusters arise from barcode sequencing errors. RESULTS: We present RUMINA, a Rust-based pipeline for UMI-aware deduplication and error correction, optimized for both amplicon and shotgun sequencing. RUMINA supports multiple UMI cluster strategies, alongside majority-rule read selection independent of mapping quality, as well as discrete handling of 1-2 read clusters, paired-end merging, and read-length stratification. Benchmarking using simulated HIV population sequencing data and real-world iCLIP and TCR datasets showed that RUMINA improves ultra-low frequency SNV detection (0.01%-1%), reduces false positives, enhances reproducibility, and processes sequencing data up to 10-fold faster than existing tools. By integrating UMI- and sequence-level correction in a high-performance framework, RUMINA offers a fast, scalable, and robust solution for UMI-enabled sequencing workflows. AVAILABILITY AND IMPLEMENTATION: RUMINA is implemented in Rust and distributed as open-source code and precompiled binaries. Source code and installation instructions are available at https://github.com/greninger-lab/rumina. Documentation associated with this manuscript is available at https://github.com/greninger-lab/rumina_paper.

High-Throughput Nucleotide Sequencing

An Application of Iterative Health Economic Evaluation: An Update on the Early Cost Effectiveness of Whole-Genome Sequencing in Advanced Non-small-Cell Lung Cancer.

OBJECTIVE: Whole genome sequencing (WGS) can identify more druggable targets than the standard of care (SoC) panels, however, its health effects and costs are highly uncertain. Given the rapidly evolving treatment landscape and pricing, an iterative approach is crucial to continuously reassess evidence and adapt economic models. Our objective was to update a previously developed economic model for WGS. METHODS: We used a structured approach to identify and report model elements requiring updates, based on established tools and methodological guidance, and applied it to the probabilistic decision model by Simons et al.(2021), which compared SoC, WGS, and SoC followed by WGS in patients with inoperable stage IIIB, C/IV NSCLC in the Dutch setting. RESULTS: Updates included a new treatment (sotorasib), revised drug and diagnostic costs, and adherence to the latest guidelines. Drug and WGS diagnostics costs fell by 8% and 26%, respectively. SoC diagnostic prices increased by 17%. We explored the impact of the prevalence of druggable targets, effectiveness of off-label treatments, (academic-specific) diagnostic costs, and price negotiations. The ICER of WGS versus SoC decreased from &#x20ac;737,197 to &#x20ac;419,053/QALY. WGS would become cost-effective if diagnostic costs descended from &#x20ac;2,180 to &#x20ac;1,246 or if additional druggable targets were identified in &#x2265;3.3% of patients. CONCLUSION: Our structured approach effectively identified items in the original analysis requiring updates and provides a foundation for further developing a checklist to guide iterative HTA. Continued monitoring and assessment of new treatment options, the dynamic diagnostics and costs throughout the life-cycle remain necessary to determine when WGS can be considered cost-effective.

NSCLC

First-line drug-resistant tuberculosis among children under 15 years in Ethiopia: insights from phenotypic and whole-genome sequencing approaches.

BACKGROUND: Childhood drug-resistant tuberculosis is often underdiagnosed and inadequately characterized due to the paucibacillary nature of the disease. This study aimed to assess resistance to first-line anti-tuberculosis drugs in children using phenotypic drug susceptibility testing and whole-genome sequencing. METHODS: A retrospective-prospective study was conducted on culture-confirmed childhood tuberculosis cases in Ethiopia (2017&#x2013;2023). Phenotypic drug susceptibility testing was performed on 110 Mycobacterium tuberculosis complex isolates. Whole-genome sequencing was completed for 85 of these isolates, which were analyzed using the TB-Profiler and MTBSeq pipelines. We assessed the sensitivity, specificity, predictive values, and kappa agreement of whole-genome sequencing compared with phenotypic drug susceptibility testing. RESULTS: Phenotypic resistance to at least one first-line anti-TB drug was observed in 26/110 (23.6%) of the examined isolates, with isoniazid resistance being the most frequent, 23/110 (20.9%), followed by rifampicin resistance, 18/110 (16.4%). TB-Profiler showed almost perfect agreement with phenotypic drug susceptibility testing for rifampicin (sensitivity 94.4%, kappa&#x2009;=&#x2009;0.96) and isoniazid (sensitivity 91.3%, kappa&#x2009;=&#x2009;0.91), whereas MTBSeq showed slightly lower performance. Both pipelines demonstrated moderate to weak agreement with phenotypic drug susceptibility testing for detecting resistance to ethambutol, pyrazinamide, and streptomycin. The most frequently observed resistance mutations among phenotypically resistant isolates were rpoB (Ser450Leu), katG (S315Thr), embB (Met306Ile), and pncA (C-11&#xa0;A&#x2009;>&#x2009;G) for rifampicin, isoniazid, ethambutol, and pyrazinamide, respectively. Discrepancies between genotypic and phenotypic drug susceptibility testing were observed across all first-line anti-TB drug-resistant isolates, particularly for ethambutol and pyrazinamide. CONCLUSION: We found a high prevalence of isoniazid resistance, along with rifampicin resistance, underscoring the need for early detection in vulnerable groups. Whole-genome sequencing showed good accuracy for these drugs, with TB-Profiler performing best. CLINICAL TRIAL NUMBER: Not applicable.

Humans

Whole-genome sequencing, as a powerful diagnostic tool in hearing loss, reveals novel variants in PTPRQ missed by whole-exome sequencing.

BACKGROUND/OBJECTIVES: Hearing loss (HL) is one of the most common congenital disorders, affecting 1-2 in 1,000 newborns. Modern genetic diagnostics using large gene panels and/or whole exome analysis (WES) can identify disease-causing mutations in 25-50&#xa0;% of patients, with higher solve rates in individuals with earlier onset. RESULTS: Here, we used whole-genome sequencing (WGS) to reanalyze 14 index patients/families who remained without genetic diagnosis by WES. We were able to identify the genetic cause of HL in 6 families (43&#xa0;%). Two families were diagnosed with DFNB84A caused by compound heterozygous recessive mutations in PTPRQ. Three of the four underlying variants, including a structural variant, a deep intronic variant, and a splice variant, escaped detection by WES. Minigene assays confirmed the pathogenicity of the intronic and the splice variants. In addition, we used protein 3D structure prediction and rigid ligand docking to study the pathogenicity of variants that escape nonsense-mediated decay. CONCLUSION: In our study, we present four novel variants in PTPRQ, three of which were detected only by WGS. To our knowledge, we report here the first pathogenic deep intronic PTPRQ variant causing HL. Our results suggest that the mutational spectrum of PTPRQ is not well covered by standard WES and that PTPRQ-associated hearing loss may be more frequent than previously thought. WGS provides an additional layer of information in the diagnostics of HL.

Humans