PubMed HealthSearch

SEARCH · PubMed Health

Results for “genome coverage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Development of a low-coverage whole genome sequencing screen for apomixis using a diverse set of Malus germplasm.

In the past decade, plant biologists have made several major discoveries pertaining to the genetic basis of apomixis (clonal propagation by seed) that have shown promise in preserving high-value hybrid rice and sorghum genotypes. This progress was made possible by foundational gene discovery efforts in model species and natural apomicts, but pleiotropic obstacles still limit its broad agricultural adoption, especially in eudicots. Thus, it follows that investigations of novel apomicts should lead to the development of new molecular tools for plant breeding. The two most common ways to identify clonal seed production are flow-cytometry seed screens and genome sequencing to compare the DNA sequences of the maternal parent and progeny, traditionally using low-throughput markers. While flow-cytometry has been the dominant method for more than two decades, it provides indirect information on the genetics of a resulting embryo and can be ineffective in certain species. Here we developed a method using short-read whole-genome sequencing at moderately low coverage (averaging 3X and 6X) to screen diverse Malus genotypes maintained in a USDA germplasm collection for clonal seed production. In total, we sequenced 55 genotypes, 1,216 of their embryos, and identified 17 previously undescribed apomictic genotypes. Several more were detected with the flow cytometry seed screen, which helped resolve certain types of reproduction and sources of noise in low-coverage datasets. This low-pass screening-by-sequencing method is a relatively low-cost, rapid method for detecting apomictic genotypes in diverse plant germplasm and when used thoughtfully in conjunction with flow cytometry, provides a new way to visualize the genetic outcomes of sexual and asexual reproduction in plants.

Apomixis

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 × homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering ∼18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals

Ancestry, admixture, and pathogens in contemporaneous Neolithic farmers and foragers on the Island of Gotland.

Two archaeological cultural complexes; the Neolithic Funnelbeaker culture (FBC) and the Pitted ware culture (PWC), coexisted on Gotland for over 500 years, between ~3300 and 2800 calBCE. The ancestry of the FBC farmers and PWC marine foragers largely aligns with European Neolithic Farmers and European Mesolithic foragers, respectively, but the direct interactions between the groups on Gotland is not understood. We present a Middle Neolithic (MN) high-coverage genome and a Late Neolithic (LN) low-coverage genome from the Ansarve FBC dolmen. We investigate ancestry, admixture, and pathogens among these MN farmers (n = 6), foragers (n = 19), and the LN individual. We find that recent gene-flow between farmers and foragers could have taken place, although most gene-flow happened prior to their coexistence on the island. We also find evidence of different Yersinia pestis strains in the three cultural groups, showing that the pestis was widespread among groups with different subsistence strategies.

Humans

LYCEUM: learning to call copy number variants on low-coverage ancient genomes.

MOTIVATION: Copy number variants (CNVs) are pivotal in driving phenotypic variation that facilitates species adaptation. They are significant contributors to various disorders, making ancient genomes crucial for uncovering the genetic origins of disease susceptibility across populations. However, detecting CNVs in ancient DNA (aDNA) samples poses substantial challenges due to several factors: (i) aDNA is often highly degraded; (ii) contamination from microbial DNA and DNA from closely related species introduces additional noise into sequencing data; and finally, (iii) the typically low-coverage of aDNA renders accurate CNV detection particularly difficult. Conventional CNV calling algorithms, which are optimized for high-coverage read-depth signals, underperform under such conditions. RESULTS: To address these limitations, we introduce LYCEUM, the first machine learning-based CNV caller for aDNA. To overcome challenges related to data quality and scarcity, we employ a two-step training strategy. First, the model is pre-trained on whole genome sequencing data from the 1000 Genomes Project, teaching it CNV-calling capabilities similar to conventional methods. Next, the model is fine-tuned using high-confidence CNV calls derived from only a few existing high-coverage aDNA samples. During this stage, the model adapts to making CNV calls based on the downsampled read depth signals of the same aDNA samples. LYCEUM achieves accurate detection of CNVs even in typically low-coverage ancient genomes. We also observe that the segmental deletion calls made by LYCEUM show correlation with the demographic history of the samples and exhibit patterns of negative selection inline with natural selection. AVAILABILITY AND IMPLEMENTATION: LYCEUM is available at https://github.com/ciceklab/LYCEUM.

DNA Copy Number Variations

Wastewater Sequencing Reveals Persistent Circulation and Rising Prevalence of Several Oncogenic Viruses Across Texas.

BACKGROUND: Oncogenic viruses cause high-risk cancers in humans and are responsible for nearly 20% of all cancer cases worldwide. Currently, very limited data exists in the realm of wastewater-based viral epidemiology (WBE) of cancer-causing viruses, with existing studies using targeted approaches (i.e PCR-based approaches) which lack scalability. Our study aims to carry out WBE with hybrid-capture probes to detect and track multiple oncogenic viruses simultaneously in wastewater across Texas, USA, overcoming the drawbacks associated with targeted approaches. METHODS: Here, we used a hybrid-capture approach to detect, filter and sequence oncogenic virus signals from wastewater samples collected over a duration of three years, from May 2022 to May 2025. Once viral reads were sequenced, we utilized established computational tools to characterize reads into their respective virus of origin. Next, viral abundances of each characterized oncogenic virus were tracked over time and read coverage across their genomes was measured using read mapping techniques. FINDINGS: We detected six known oncogenic viruses, along with three suspected oncogenic viruses across all sampling locations within Texas. Over three years, viral abundance gradually increased, with distinct peaks and dips over the summer and winter months. The prevalence of high-risk viruses such as HPV and EBV rose sharply, with increases in abundance observed post-2024. We also obtained nearly 100% genome coverage with viral reads captured using a hybrid-capture technique for almost all oncogenic viruses and their types. INTERPRETATIONS: Our study shows that a hybrid-capture method can efficiently overcome the challenges faced with using targeted approaches for WBE. Using this method, we get broader read coverage, coupled with concurrent and consistent real-time tracking dynamics of multiple oncogenic viruses. Our findings also emphasize the persistent circulation and rising prevalence of high-risk cancer-causing viruses, underscoring the need for sustained public health interventions to protect communities and assess viral prevalence in high-risk populations. FUNDING: This work was supported by S.B. 1780, 87th Legislature, 2021 Reg. Sess. (Texas 2021), the Baylor College of Medicine and the Alkek Foundation Seed Funds.

Journal Article

Nanopore-based sequencing of active DNA replication reveals key principles of metazoan replication fork progression, origin and termination sites.

Balancing replication fork progression and origin usage is essential to maintain genome stability, but measuring replication fork progression rates and origin usage throughout the genome has been challenging. Here, we use nanopore sequencing combined with DNAscent to measure replication fork progression together with origin and termination site usage with single-molecule precision throughout the Drosophila genome with nearly full genome coverage. We find that replication fork progression rates are not uniform throughout the genome. Rather, fork progression is slowest in euchromatin, and this is not correlated with active transcription. Replication origins are also influenced by chromatin, but the exact position of initiation is highly variable and are often several kilobases away from ORC binding sites. Termination sites lack any chromatin or sequence motifs and appear nearly random throughout the genome. By measuring DNA replication dynamics at near full genome coverage, our work reveals key principles of metazoan replication dynamics.

Journal Article

Reflective Evaluation of Next-Generation Sequencing Data during Early Phase Detection of the Delta Variant.

During the SARS-CoV-2 pandemic, next-generation sequencing (NGS) technologies like the Ion Torrent S5 and Illumina MiSeq, alongside advanced software, improved genomic surveillance in South Africa. This study analysed anonymized samples from the Eastern Cape using Genome Detective and NextClade, showing Ion Torrent S5 and Illumina MiSeq success rates of 96% and 94%, respectively. The study focused on genomic coverage (above 80%) and mutation detection (below 100), with the Ion Torrent S5 achieving 99% coverage compared to Illumina MiSeq's 80%, likely due to different primers used in amplification. The Ion Torrent S5 was more effective in sequencing varied viral loads, whereas Illumina MiSeq had difficulties with lower loads. Both platforms were adept at identifying clades, successfully differentiating between Beta (<45%) and Delta variants (<30%), despite minor discrepancies in assignments due to Illumina MiSeq's lower coverage, leading to a failure rate of up to 6%. Manual library preparation showed similar sample processing and clade identification capabilities for both platforms. However, differences in sequencing duration (3.5 vs. 36 hours), automation level, genomic coverage (80% vs. 99%), and viral load compatibility were noted, highlighting each platform's unique advantages and challenges in SARS-CoV-2 genomic surveillance. In conclusion, the Illumina MiSeq and Ion Torrent S5 platforms are both efficacious in executing whole-genome sequencing (WGS) via amplicons, facilitating precise, accurate, and high-throughput examinations of SARS-CoV-2 viral genomes. However, it is important to note the existence of disparities in the quality of data produced by each platform. Each system offers unique benefits and limitations, rendering them viable choices for the genomic surveillance of SARS-CoV-2.

Illumina MiSeq

Primer design through submodular function estimation.

MOTIVATION: Multiplex PCR-based enrichment is widely used in viral genome sequencing and pathogen surveillance. However, designing large sets of primers that maximize genome coverage while minimizing primer-primer interactions remains a major computational challenge. Existing methods such as SADDLE and Olivar use heuristics to optimize a Badness score for primer dimers but lack theoretical guarantees on solution quality. RESULTS: We introduce PRISM, a new framework that formulates multiplex primer design as a constrained submodular maximization problem. Our method defines an objective that balances genome coverage and dimer risk, and applies a local search algorithm with a constant-factor approximation guarantee. Evaluations on viral genome datasets demonstrate that PRISM consistently achieves lower Badness scores compared to PrimalScheme, Olivar, and primerJinn. These results highlight the scalability and theoretical rigor of submodular optimization in primer design. AVAILABILITY: PRISM is open-source and available at https://github.com/yhhan19/PRISM-new. The experimental data, scripts, and results used in this paper are archived on Figshare at https://doi.org/10.6084/m9.figshare.32806499.

Algorithms

Development and evaluation of an ARTIC-based amplicon sequencing assay for whole-genome characterization of respiratory syncytial virus.

Respiratory syncytial virus (RSV), a ~15.2 kb negative-sense RNA virus, causes acute respiratory infections in infants and older adults. Its two subtypes, RSV-A and RSV-B, evolve rapidly, making ongoing monitoring of circulating strains essential. The Georgia Public Health Laboratory (GPHL) developed and evaluated an amplicon-based whole-genome sequencing (WGS) assay for RSV surveillance. A total of 214 de-identified remnant clinical specimens (102 RSV-A and 112 RSV-B) with RT-PCR Cq values <31 were included. RSV genomes were amplified using ARTIC-style and custom primer sets, with the ARTIC set showing superior performance. Libraries were prepared using a modified Illumina COVIDSeq protocol, sequenced on NextSeq 1000/2000 instruments, and analyzed using the GPHL-RSV-PIPE bioinformatics pipeline. Among genomes meeting validation criteria, sequencing depth was slightly higher for RSV-A (median 53,433&#xd7;; mean 51,076&#xd7;) than RSV-B (median 49,699&#xd7;; mean 46,945&#xd7;), whereas genomic coverage was slightly lower for RSV-A (median 97.5%; mean 96.6%) than RSV-B (median 98.3%; mean 97.6%). Predominant lineages were A.D.3.1 and A.D.5.2 for RSV-A and B.D.E.1 for RSV-B. For RSV-A, the assay showed 92.8% accuracy, 96.2% sensitivity, 87.2% specificity, 92.6% positive predictive value, and 93.2% negative predictive value. Intra- and inter-run precision assessed using 16 and 53-57 genomes, respectively, showed nearly 100% consensus genome identity with 0-5 nucleotide differences. Specificity testing of 31 non-RSV specimens produced no false-positive detections. Limits of detection were 4.4 TCID50/mL for RSV-A and 18.6 TCID50/mL for RSV-B. These results demonstrate that the ARTIC-based RSV WGS assay enables near real-time surveillance and strengthens data-driven public health responses to future outbreaks.IMPORTANCERSV, with two major subtypes, RSV-A and RSV-B, causes acute respiratory infections that can be severe in infants under 6 months and older adults. Current RSV surveillance at the GPHL relies on the Thermo Fisher TaqMan Gene Expression Capillary assay, which detects and subtypes RSV but lacks resolution for lineage classification and identification of emerging variants. To address this critical gap, GPHL developed and evaluated an amplicon-based WGS assay using 214 de-identified RSV clinical specimens. Genomes were amplified using ARTIC-style and custom-primer sets, with ARTIC primers showing superior performance. The assay demonstrated strong sequencing depth, genomic coverage, specificity, repeatability, reproducibility, and low limits of detection. RSV lineages were accurately determined based on genetic variation. These results establish that the ARTIC-based WGS assay enables near real-time genomic surveillance, supporting monitoring of circulating RSV strains and informing data-driven public health responses.

bioinformatics pipeline

Benchmarking DNA extraction protocols across use cases for culture-independent Nanopore metagenomics.

Oxford Nanopore Technologies (ONT) sequencing offers several advantages for metagenomics, including long reads, rapid turnaround, low upfront cost, scalability and portability. However, for ONT metagenomics, DNA yield, quality and integrity are important considerations when selecting an extraction method. Many metagenomic extraction methods use harsh lysis conditions to extract a wide range of species and provide an accurate community composition, but these conditions can compromise DNA fragment length. Therefore, extraction methods for ONT metagenomics must balance DNA shearing and recovery with representative community lysis. We systematically evaluated DNA extraction methods for ONT metagenomic sequencing using a use case-oriented framework. Among nearly 50 extraction methods screened, 7 were selected for detailed comparison based on suitability for metagenomics, variation in methodology, availability, cost and processing time: Norgen BioTek Corp's Stool DNA Isolation (NG), Zymo Research's ZymoBIOMICS Quick-DNA HMW MagBead (ZMG), Qiagen's DNeasy Blood and Tissue (QBT), Macherey-Nagel's NucleoMag DNA Microbiome (MN), Zymo Research's ZymoBIOMICS DNA Mini Prep (ZMI), Qiagen's DNeasy PowerSoil/QIAamp PowerFecal Pro (PS) and Qiagen's QIAamp Fast DNA Stool Mini (QIA). Methods were tested using Zymo Research's ZymoBIOMICS Microbial Community Standard (MCS), a matrix-free mock community with known composition. DNA extracts were sequenced on an ONT PromethION using the Rapid Barcoding Kit, except QIA due to insufficient DNA yield. Metrics for the method, DNA extracts, sequencing and genomes were evaluated, revealing trade-offs between methods. The two magnetic bead methods, MN and ZMG, produced the highest mean read length N50 values (13.9 and 16.5&#x2009;kb, respectively) but showed apparent community compositions skewed towards Gram-negative bacteria. In contrast, ZMI and PS maintained a community composition close to expected, with reduced mean read length N50 values (4.5 vs. 7.5&#x2009;kb). Performance across various metrics is presented in the context of the following use cases: maximizing genome coverage and assembly completeness, preserving composition accuracy, targeting specific species and limiting required resources (equipment, time or budget). The metrics and use case considerations presented offer practical guidance for informed selection of DNA extraction methods for ONT metagenomics. For accurate community composition, ZMI or PS are recommended, while PS and ZMG perform best at maximizing genome coverage and assembly completeness. NG and QBT may be the most economical options, though performance trade-offs were observed. Finally, PS may be the preferred method for time-sensitive diagnostic or field applications.

Metagenomics

IBDV-SSA, a novel molecular approach for the recovery of infectious bursal disease virus whole genomes from FTA cards.

Infectious bursal disease (IBD), a highly contagious viral disease in young chickens, poses significant economic losses due to high mortality and immunosuppression. While IBD virus (IBDV) virulence is influenced by multiple genes, whole-genome sequencing (WGS) of IBDV is crucial for defining the strain pathotype and clinical profile. Flinders Technology Associates (FTA) cards are convenient for field sample collection, but their filter paper matrix can hinder nucleic acid recovery, impacting sequencing efficiency. This study evaluated two enrichment strategies, single primer amplification (SPA) and IBDV segment-specific amplification (SSA), coupled with short-read (Illumina) and long-read (Oxford Nanopore Technologies, ONT) sequencing platforms, to optimize IBDV whole-genome recovery from FTA cards. Illumina sequencing produced comparable raw read counts for both methods, yet IBDV-SSA samples achieved significantly higher genome mapping rates (76%) than IBDV-SPA (12%). Genome coverage analysis revealed that IBDV-SSA provided uniform read distribution across both genomic segments, ensuring complete coverage, while IBDV-SPA exhibited significant bias, with most reads mapping to segment B, and limited coverage of segment A. Importantly, IBDV-SSA also proved compatible with ONT long-read sequencing, providing complete genome coverage. Notably, IBDV-SSA coupled with short-read sequencing successfully characterized coinfections in two samples. This optimized approach using IBDV-SSA enables efficient and comprehensive WGS of IBDV from FTA cards, facilitating strain characterization, virulence prediction, and epidemiological investigations.IMPORTANCEThis research tackles a significant problem for poultry farmers: a virus called infectious bursal disease virus (IBDV) that harms young chickens, causing high death rates and economic losses. To fight it effectively, scientists need to analyze its complete genetic makeup. Traditionally, collecting and preserving IBDV field samples was challenging. Flinders Technology Associates (FTA) cards have simplified this process, but getting usable genetic material from them has been difficult. This study introduces a new genome enrichment method, IBDV segment-specific amplification (IBDV-SSA), which successfully allows for IBDV complete genome recovery from FTA cards. By using this improved approach, scientists can accurately identify virus strains, assess how harmful they are, and monitor their spread. This, in turn, helps to improve vaccines and protect flocks. IBDV-SSA is a powerful tool for outbreak surveillance, supporting the poultry industry and ensuring a stable food supply.

Infectious bursal disease virus

Genomic Analysis of Circulating Tumor Cells at the Single-Cell Level.

Circulating tumor cells (CTCs) have a great potential for noninvasive diagnosis and real-time monitoring of cancer. A comprehensive evaluation of four whole genome amplification (WGA)/next-generation sequencing workflows for genomic analysis of single CTCs, including PCR-based (GenomePlex and Ampli1), multiple displacement amplification (Repli-g), and hybrid PCR- and multiple displacement amplification-based [multiple annealing and loop-based amplification cycling (MALBAC)] is reported herein. To demonstrate clinical utilities, copy number variations (CNVs) in single CTCs isolated from four patients with squamous non-small-cell lung cancer were profiled. Results indicate that MALBAC and Repli-g WGA have significantly broader genomic coverage compared with GenomePlex and Ampli1. Furthermore, MALBAC coupled with low-pass whole genome sequencing has better coverage breadth, uniformity, and reproducibility and is superior to Repli-g for genome-wide CNV profiling and detecting focal oncogenic amplifications. For mutation analysis, none of the WGA methods were found to achieve sufficient sensitivity and specificity by whole exome sequencing. Finally, profiling of single CTCs from patients with non-small-cell lung cancer revealed potentially clinically relevant CNVs. In conclusion, MALBAC WGA coupled with low-pass whole genome sequencing is a robust workflow for genome-wide CNV profiling at single-cell level and has great potential to be applied in clinical investigations. Nevertheless, data suggest that none of the evaluated single-cell sequencing workflows can reach sufficient sensitivity or specificity for mutation detection required for clinical applications.

Carcinoma, Non-Small-Cell Lung

Use of high coverage reference libraries of Drosophila melanogaster for relational data analysis. A step towards mapping and sequencing of the genome.

Three differently made, primary Drosophila cosmid libraries of 16-fold genome coverage have been generated. Also, a jumping library has been created by a new method that takes advantage of methylation differences between genomic DNA and vector. Thirdly, two cDNA libraries have been picked. All these libraries have been arrayed on high-density in situ filters, each containing 9216 clones. As a reference system, such filters are distributed and identified clones are provided. Single-copy probes have identified on average 1.4 cosmids per genome equivalent. Together with cytogenetically mapped yeast artificial chromosomes, the libraries are also being used for physically mapping the genome, mainly by oligonucleotide fingerprinting and pool hybridizations. cDNA clones are further examined by a partial sequencing analysis by oligomer hybridization.

Animals

JG2: an updated version of the Japanese population-specific reference genome.

Here we present the construction of JG2, an updated population-specific reference genome for the Japanese population. Utilizing data from three individuals previously used in the construction of JG1, several methodologies were employed to enhance genomic coverage and assembly quality. Hi-C sequencing technology facilitated phase-aware assembly, generating two haploid assemblies per individual and enabling improved representation of genetic variation. A meta-assembly strategy and a majority decision approach further refined assembly quality by combining the best sequences from multiple assemblies and minimizing the inclusion of rare variants. The resulting JG2 genome comprises chromosome-level sequences, mitochondrial chromosomes and unplaced scaffolds, offering more comprehensive coverage of the Japanese genome. Comparative analyses with other reference genomes demonstrated the accuracy and representativeness of JG2, highlighting its utility for genetic research involving the Japanese population. Overall, by adopting the phased assembly technique, JG2 represents a substantial advancement over the collapsed assembly-based JG1, with improvements including a greater number of identified variants (3,115,695 variants, of which 298,644 had an allele frequency (AF) of 1.0 in the 3.5KJPNv2 AF panel) and a higher N50 value (152,668,378&#x2009;bp). These enhancements provide researchers with a more precise and comprehensive resource for understanding the genetic landscape of the Japanese population. The sequences and annotations are available on the jMorp website ( https://jmorp.megabank.tohoku.ac.jp/ ).

Journal Article

The offonome reveals on and off states of gene expression near the detection limit of RNA-seq.

RNA-seq, widely used for gene expression profiling, provides nucleotide level genome coverage and summary gene expression values. Generally, low-expressed genes are ignored due to their unfavorable signal-to-noise ratio, however, these genes may offer crucial information, such as detecting rare cells in bulk tissues. In this study, we applied an approach that transforms the expression levels of low-expressed genes into a robust dichotomized on/off state by leveraging similarities in transcript coverage shape. Applied to three human cancer cohorts from the Cancer Genome Atlas (TCGA), chosen based on tissue morphology and anatomic site, we identified genes, the "offonome" near the detection limit, consistently or occasionally off across samples. Genes in the offonome spectrum proved useful for supervised and unsupervised applications, including characterizing oncogenic pathways, and identifying rare populations of cells in bulk tissue. Interrogating the offonome is relevant to bulk tumor analyses like TCGA, potentially expediting gene investigation in low-input situations like single cell RNA-seq.

Humans

ExoShorkie: predicting RNA-seq coverage of exogenous genomes in yeast by transfer learning.

MOTIVATION: Predicting the RNA-seq coverage of native and exogenous sequences is central to many molecular- and synthetic-biology applications. Substantial progress has been made in developing methods to predict the RNA-seq coverage of native genomic sequences, with the recently developed Shorkie achieving state-of-the-art performance in yeast. However, prediction performance of these methods over exogenous DNA is still unknown. Recent studies measured RNA-seq coverage of large exogenous genomes in yeast, providing a unique opportunity to train machine-learning models on a large exogenous sequence space and to improve both prediction performance and our understanding of regulatory mechanisms. RESULTS: We introduce ExoShorkie, a method we developed by extending Shorkie through transfer learning across multiple exogenous RNA-seq datasets. We demonstrate that ExoShorkie significantly improves prediction performance on held-out exogenous genomes and outperforms both a native-genome-trained Shorkie baseline and Yorzoi, the only competing method for predicting exogenous RNA-seq coverage in yeast, in cross-validation and in leave-one-genome-out evaluations. Furthermore, through interpretability analyses we reveal biologically meaningful regulatory motifs and distinct regulatory rules in exogenous genomes in yeast, providing new insights into transcriptional regulation. AVAILABILITY AND IMPLEMENTATION: ExoShorkie is available at https://github.com/OrensteinLab/ExoShorkie.

Genome, Fungal

Whole genome sequence of a superbug-Escherichia coli strain KAB-AI-497 isolated from the vagina of a 20 year old pregnant woman with premature rupture of membrane (PROM) in a resource limited setting, Kabale Regional Referral Hospital, in Uganda.

OBJECTIVES: The objective of the study is to sequence the whole genome of multidrug resistant E. coli strain KAB-AI-497 that causes bacterial vaginosis and implicated in premature rupture of membrane in pregnant woman. DATA DESCRIPTION: The DNA of the E. coli strain KAB-AI-497 was extracted using the MagAttract HMW DNA Kit, and the extracted DNA was sequenced using an MGI DNBSEQ G99ARS platform. FastQC was used to perform quality control analysis and the reads were trimmed by Trimmomatic. De novo genome assembly was performed by SPAdes and it resulted to a draft assembled genome that has 5.1 Mb genome size, 153 contigs, and 50.5% GC content. Quality analysis of the assembled genome revealed it has 98.46% completeness and 0.97% contamination. The closest E. coli strain to this strain KAB-AI-497 in terms of similarity was Escherichia coli SMS-3-5 with an average nucleotide identity of 98.43% and genome coverage of 86.19%, which confirmed the species level identity of the strain. The assembled genome was annotated using the NCBI Prokaryotic Genome Annotation Pipeline which identified 4,726 protein coding genes in the strain genome. Furthermore, the annotation revealed the genome has resistant genes responsible for resistance against many antibiotic classes such as tetracycline, fluoroquinolone, and penicillin.

Escherichia coli

MRDtarget: A heuristic Gaussian approach for optimizing targeted capture regions to enhance Minimal Residual Disease detection.

Molecular residual disease (MRD) detection, initially developed for hematologic malignancies, has become a critical biomarker for monitoring solid tumors. MRD detection primarily relies on circulating tumor DNA (ctDNA) analysis using next-generation sequencing, offering high sensitivity and broad genomic coverage. However, challenges remain in designing cost-effective panels that maximize mutation detection while maintaining biological relevance. Fixed panels often lack sufficient patient-specific mutation coverage, while WES-based personalized MRD assays, despite their high sensitivity, are costly and less accessible. We developed a tumor comprehensive genomic profiling (CGP)-informed personalized MRD assay to detect tumor-derived mutations, which allowed us to design patient-specific personalized panels and meanwhile, provide a cost-effective alternative to whole exome sequencing (WES). To address these limitations, we developed MRDtarget, a heuristic multivariate Gaussian model-based targeted capture region selection method. By expanding beyond traditional hotspot regions, MRDtarget optimizes variant tracking for MRD detection, significantly improving sensitivity. Using a Bayesian inference-based heuristic approach, MRDtarget integrates multi-feature informativeness rates to identify optimal genomic regions for capture. Experimental results demonstrate that MRDtarget enables the detection of more variants per patient. This study underscores the importance of rational panel design to improve MRD sensitivity and provides a novel approach to enhance precision diagnostics and treatment for solid tumor patients.

Humans