PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genome Sequence Archive”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The GSA Family in 2025: A Broadened Sharing Platform for Multi-omics and Multimodal Data.

The Genome Sequence Archive family (GSA family) provides a comprehensive suite of database resources for archiving, retrieving, and sharing multi-omics data for the global academic and industrial communities. It currently comprises four distinct database members: the Genome Sequence Archive (GSA, https://ngdc.cncb.ac.cn/gsa), the Genome Sequence Archive for Human (GSA-Human, https://ngdc.cncb.ac.cn/gsa-human), the Open Archive for Miscellaneous Data (OMIX, https://ngdc.cncb.ac.cn/omix), and the Open Biomedical Imaging Archive (OBIA, https://ngdc.cncb.ac.cn/obia). Compared to its 2021 version, the GSA family has expanded significantly by introducing a new repository, the OBIA, and by comprehensively upgrading the existing databases. Notable enhancements to the existing members include broadening the range of accepted data types, strengthening quality control systems, improving the data retrieval system, and refining data-sharing management mechanisms.

Humans

Genomic population structure, antimicrobial susceptibility, and clinical features of Mycobacterium xenopi isolates, Frankfurt, Germany, 1995-2020.

Mycobacterium xenopi causes non-tuberculous mycobacterial pulmonary disease (NTM-PD) that is difficult to treat. However, data on the genomic population structure, antimicrobial susceptibility, and the clinical significance of this pathogen remain scarce. We analyzed 76 clinical M. xenopi isolates from 70 patients collected between 1995 and 2020 in Frankfurt am Main, Germany. All isolates underwent phenotypic drug susceptibility testing and whole-genome sequencing. Cluster analysis, including isolates from this study and all hitherto available high-quality M. xenopi genome data sets in the Sequence Read Archive (n = 11), was performed by core genome multilocus sequence typing. In our cohort, only 26.5% of patients met criteria for clinically relevant NTM-PD. Phylogenetic analysis identified three large hospital-associated clusters (≤10 allelic difference), each involving between 7 and 20 patients and persisting for over 18 years, suggesting prolonged transmission chains or a common environmental source. We also defined three major clades (≤50 allelic difference), two of which contained isolates from the United Kingdom. Clofazimine and guideline-recommended antimycobacterial agents showed good in vitro efficacy, except rifampicin, with 23.6% resistance. This study represents a major expansion of M. xenopi genomic resources and provides insights into the genomic population structure, phenotypic susceptibility, and clinical characteristics of M. xenopi. Guideline-recommended antimycobacterials show good in vitro activity, while clofazimine may be a valuable addition to M. xenopi therapy. The identified clusters underscore the need for further investigation into transmission dynamics and globally successful clones.IMPORTANCEMycobacterium xenopi is an increasingly recognized opportunistic lung pathogen that is difficult to treat. Infections often occur in patients with pre-existing health conditions and can present substantial diagnostic and therapeutic challenges. A deeper understanding of its genetic diversity and resistance mechanisms is essential for optimal patient management and for clarifying potential transmission routes. By analyzing 76 whole-genome sequences together with detailed clinical information and phenotypic drug-susceptibility data, this study substantially expands the available genomic repertoire for M. xenopi. While clinical relevance was limited in our cohort, most guideline-recommended antimicrobial agents showed good efficacy in vitro. The detection of closely related strains might point toward a common environmental source of infection. These findings highlight the need for continued surveillance and provide a comprehensive foundation that supports more accurate monitoring, improved understanding of disease behavior, and future investigations into M. xenopi pathogenicity.

Humans

Embryophyte-wide detection of natural Agrobacterium-mediated horizontal gene transfer reveals an ancient role for mini T-DNAs.

Agrobacterium transfers DNA into plant cells, leading to tumors, hairy roots (HR), and natural genetically modified organisms (nGMOs). Transferred DNAs (T-DNAs) from agrobacteria and T-DNA-derived cellular T-DNAs (cT-DNAs) from nGMOs vary considerably and may carry up to 15 different genes. Among these, opine synthase (ops) genes encode the synthesis of opines used as nutrients by the agrobacteria. Earlier studies predicted large numbers of naturally transformed plant species, but only few have been identified and studied so far. We therefore developed a general method to detect cT-DNAs in all publicly available whole genome sequences (WGS) and Sequence Read Archive (SRA) data from land plants. To avoid false positives, we only retained DNA sequences coding for T-DNA proteins. A total of 2614 nGMO species were identified, most are eudicots. However, cT-DNAs were also found in 82 mosses and 75 ferns, showing that Agrobacterium can also generate natural transformants among the early land plants. Analysis of 149 cT-DNA maps revealed different types of T-DNAs. Most notably, these included small T-DNAs (mini T-DNAs) with a single opine synthase gene. Mini T-DNAs are not expected to induce tumors or HRs. The predominance of mini cT-DNAs in mosses and ferns, and the presence of more complex cT-DNAs in spermatophytes, indicate that mini T-DNAs represent the earliest types of T-DNA. Our study also detected unusual T-DNA integration patterns, with multiple copies spread out over several hundreds of kilobases.

DNA, Bacterial

Universal versus targeted chlorhexidine and mupirocin decolonisation and clinical and molecular epidemiology of Staphylococcus epidermidis bloodstream infections in patients in intensive care in Scotland, UK: a controlled time-series and longitudinal genotypic study.

BACKGROUND: There are concerns that biocide skin and mucous membrane decolonisation, which is widely used to prevent health-care-associated infections in intensive care units (ICUs), might select for multidrug-resistant pathogens. We aimed to evaluate the effects of de-escalating from universal to targeted skin and nasal decolonisation on Staphylococcus epidermidis bloodstream infections (SE-BSI). METHODS: We did a retrospective, before-after-control-impact time-series analysis and longitudinal genotypic study in two ICUs with divergent decolonisation practice in tertiary care hospitals of adjacent health boards in Scotland, UK. Participants were aged at least 16 years and admitted between July 1, 2009, and Feb 28, 2022. There were no exclusion criteria for the study. In ICU one (intervention site) universal decolonisation in all admissions was de-escalated to targeted decolonisation of meticillin-resistant Staphylococcus aureus (MRSA) carriers on Feb 1, 2019, while in ICU two (control site) targeted decolonisation was applied throughout. We collected bloodstream infection data from all causes, including clinically significant SE-BSI. Antimicrobial susceptibility testing was used to define meticillin-resistant S epidermidis (MRSE) and chlorhexidine susceptibility. We used multilocus sequence typing to identify sequence types from archived SE-BSI isolates. Whole-genome sequencing was applied to a sample from ICU one. The primary outcomes were incidence densities of all bloodstream infections, SE-BSI, and meticillin-resistant S epidermidis bloodstream infections (MRSE-BSI), and the percentage probability that SE-BSI were MRSE-BSI. The effects of de-escalation on primary outcomes were estimated by differences between the intervention and control sites, before and after de-escalation, using a before-after-control-impact time-series design. Secondary outcomes included the proportion of multidrug resistant sequence types, carriage of mobile genetic elements and genes for multidrug resistance and biofilm production. FINDINGS: Between July 1, 2009, and Feb 28, 2022, S epidermidis was identified in 334 (45%) of 735 bloodstream infections in ICU one, of which 197 occurred before the de-escalation intervention in Feb 1, 2019, and S epidermidis was identified in 167 (60%) of 278 bloodstream infections in ICU two. There was no increase in all bloodstream infection incidence coinciding with de-escalation in ICU one, whereas MRSE-BSI incidence declined significantly from 10·4 cases per 1000 occupied bed days (OBDs; 95% credible interval [CrI] 7·2-15·4) to 4·3 cases per 1000 OBDs (2·5-6·7), as did the percentage probability of MRSE (from 89·2%, 95% CrI 77·8-96·5 to 56·7%, 34·3-77·5%). No significant changes in the primary outcomes were seen in ICU two. MRSE-BSI incidence density was positively associated with chlorhexidine use, but not mupirocin use. De-escalation was associated with a reduced proportion of SE-BSI due to multidrug-resistant sequence types and reduced carriage of mobile genetic elements and genes for multidrug resistance and biofilm production, as observed by multi-locus sequence typing and whole genome sequencing. INTERPRETATION: In ICU settings with low MRSA incidence, the benefits of universal decolonisation should be balanced against the risks of selecting MRSE sequence types adapted for invasive and device-associated infection. FUNDING: National Health Service Grampian Charity.

Humans

Whole-genome sequencing of adenovirus 41 directly from wastewater using nested overlapping PCR and MinION.

Human adenovirus F41 (HAdV-F41) is one of the leading causes of children's acute gastroenteritis and was recently linked to an outbreak of severe acute hepatitis of unknown etiology among children during 2021 to 2022. While most evidence is based on clinical data, wastewater-based epidemiology offers a community-level approach to monitoring circulating strains and enhancing outbreak preparedness. In this study, we developed an overlapping amplicon-based whole-genome sequencing approach to directly detect HAdV-F41 from archived wastewater samples, using nested PCR with 13 primer sets. Archived wastewater samples were collected between 2021 and 2022 from three treatment plants in Seattle, USA. The viral load ranged from 1.2 × 103 to 8.4 × 103 genome copies per liter. The Oxford Nanopore platform was used for whole-genome sequencing. Complete or partial (>84%) HAdV-F41 genomes were recovered from wastewater samples, with mean coverage depths ranging from 10³ to 10⁵. The consensus sequences showed more than 99% similarity to reference genomes in the NCBI database. The phylogenetic analysis revealed that 2 sequences clustered within lineage 2a and 11 within lineage 2b, reflecting that at least two sub-lineages were circulating in the community at that time. Our results demonstrate that the overlapping amplicon-based whole-genome sequencing approach using the Oxford Nanopore platform reliably recovers HAdV-F41 genomes from wastewater. This method offers high-resolution genomic surveillance of circulating, clinically relevant HAdV-F41, supporting wastewater-based epidemiology as a valuable tool for detecting emerging variants and strengthening the early warning system for future disease outbreaks.IMPORTANCEHuman adenovirus F41 is a primary cause of childhood gastroenteritis and has been linked to recent outbreaks of severe acute hepatitis in children, yet community-level genomic surveillance of this virus remains limited. This study shows that wastewater can be used to recover nearly complete HAdV-F41 genomes through a targeted overlapping-amplicon sequencing strategy on the Oxford Nanopore platform. By applying this method to archived wastewater samples, we detected the simultaneous circulation of multiple viral lineages in a large city. These findings extend wastewater-based epidemiology beyond SARS-CoV-2 and emphasize its importance for monitoring clinically significant enteric viruses. The method described here offers a scalable tool for tracking viral evolution in communities and enhancing early warning systems for future outbreaks.

Wastewater

Genetic characterization of rat hepatitis E virus (Rocahepevirus ratti) in urban brown rats (Rattus norvegicus) in Helsinki, Finland.

We report complete and partial genome sequences of rat hepatitis E virus (RHEV, Rocahepevirus ratti), from archived brown rats captured in Helsinki, Finland. Phylogenetic analysis confirmed the presence of the pathogenic RHEV genotype C1 in the Helsinki region. Finnish strains clustered together with strains from South Korea and Spain. However, the polytomous topology of phylogenetic trees and the large genetic distances between spatially distinct strains suggest that RHEV has remained inadequately sampled on a global scale. Further surveillance of rocahepeviruses is needed to assess their threat to public health and to understand their diversity and evolutionary patterns.

Animals

Genomic determinants of antibiotic resistance for Helicobacter pylori treatment: a retrospective phenotypic and genotypic observational study.

BACKGROUND: Rising antimicrobial resistance of Helicobacter pylori is a public health challenge. Genomic-based susceptibility testing allows for the identification of resistance-associated mutations, complementing conventional diagnostics and advancing towards pathogen-based personalised therapies. Our study aimed to identify genes and mutations involved in antimicrobial resistance in H pylori and evaluate the extent to which these markers can be used as predictors of phenotypic resistance against clarithromycin and levofloxacin. METHODS: In this retrospective phenotypic and genotypic observational study, we included 1011 H pylori whole-genome sequences and strains of known geographical origin from the H pylori Genome Project (HpGP) collection. We performed phenotypic clarithromycin and levofloxacin susceptibility testing on a subset of 419 HpGP strains using Etest at a centralised laboratory. A genomic analysis was conducted to identify 23S rRNA and gyrA variants and build a curated catalogue of mutations associated with resistance to clarithromycin (ie, 23S rRNA 2142A→G, 2142A→C, and 2143A→G) and levofloxacin (ie, gyrA A88V or A88P, N87K or N87I, and D91G, D91N, or D91Y). Genotype-phenotype concordance was assessed to estimate sensitivity and specificity, and the curated catalogue of resistance-associated mutations was applied to the complete HpGP set. Region-specific prevalence of resistance-associated mutations was calculated for a combined dataset including the HpGP genomes and 768 whole-genome sequences retrieved from the US National Center for Biotechnology Information Sequence Read Archive repository. Associations between resistance genotypes, H pylori subpopulations, and minimum inhibitory concentrations (MICs) were tested. FINDINGS: Clarithromycin-resistant and levofloxacin-resistant HpGP strains were estimated with a sensitivity and specificity of 100%, with all confidence intervals ranging from 96% to 100%. The combined analysis (n=1779) found the highest prevalence of clarithromycin resistance in the western Pacific region (173 [51·2%] of 338 in southeast Asia and 75 [29·8%] of 252 in eastern Asia), north African region (seven [38·9%] of 18), and western Asian region (12 [31·6%] of 38), whereas the highest prevalence of levofloxacin resistance was found in south Asia (14 [51·85%] of 27), Central America (48 [38·7%] of 124), eastern Europe (four [36·4%] of 11), and southern Africa (three [33·3%] of nine). Similarly, 23S rRNA and gyrA genotypes are variable across H pylori subpopulations. MIC values changed depending on the specific mutation in 23S rRNA (mean clarithromycin MIC 24·61 mg/L [95% CI 12·27-36·96] for 2143A→G and 142·25 mg/L [95% CI 77·88-206·61] for 2142A→G) and gyrA (mean levofloxacin MIC 9·66 mg/L [95% CI 6·75-12·56] for mutations on codon 91, and 27·97 mg/L [95% CI 25·82-30·11] for mutations on codon 87). INTERPRETATION: Mutations in specific genes are reliable indicators to clarithromycin and levofloxacin resistance in H pylori, making them useful markers for the development of diagnostic assays and molecular monitoring. Our results suggest that using clarithromycin and levofloxacin empirically, without previous susceptibility testing, is unsuitable in all geographical regions covered by this study. FUNDING: Intramural Research Program of the US National Cancer Institute, the European Research Council, and the Spanish Ministry of Science and Innovation.

Helicobacter pylori

Epstein-Barr virus detection in neck metastases by polymerase chain reaction.

Cervical nodal metastasis from occult carcinomas represents a diagnostic challenge. This is a common presentation of undifferentiated nasopharyngeal carcinoma (UNPC), but metastatic carcinomas from other sites must be considered. UNPC has the distinguishing feature of a close association with Epstein-Barr virus (EBV). Since the polymerase chain reaction (PCR) can detect EBV in archival tissues, it offers significant advantages over previous methods for the detection of viral genomes. Its extreme sensitivity allows analysis of small samples from needle aspirates. Using the polymerase chain reaction to amplify EBV sequences from archival tissues, 15 of 18 NPC samples were positive for EBV. Of these 18, 14 of 14 with UNPC were positive, 1 of 2 with squamous cell carcinoma (SCC) were positive, and 10 of 2 with adenocarcinoma were positive. All 6 UNPC metastatic to lymph nodes were positive. Carcinoma metastatic to cervical nodes from 17 of 17 non-UNPC occult primaries lacked EBV. This demonstrates the utility of EBV detection by the polymerase chain reaction in the evaluation of patients with metastases to neck nodes from occult primary carcinomas in order to identify cases of UNPC.

Adult

Drug resistance mutations in HIV provirus are associated with defective proviral genomes with hypermutation.

BACKGROUND: HIV proviral sequencing overcomes the limit of plasma viral load requirement by detecting all the 'archived mutations', but the clinical relevance remains to be evaluated. METHODS: We included 25 participants with available proviral sequences (both intact and defective sequences available) and utilized the genotypic sensitivity score (GSS) to evaluate the level of resistance in their provirus and plasma virus. Defective sequences were further categorized as sequences with and without hypermutations. Personalized GSS score and total GSS score were calculated to evaluate the level of resistance to a whole panel of antiretroviral therapies and to certain antiretroviral therapy that a participant was using. The rate of sequences with drug resistance mutations (DRMs) within each sequence compartment (intact, defective and plasma viral sequences) was calculated for each participant. RESULTS: Defective proviral sequences harbored more DRMs than other sequence compartments, with a median DRM rate of 0.25 compared with intact sequences (0.0, P&#x200a;=&#x200a;0.014) and plasma sequences (0.095, P&#x200a;=&#x200a;0.30). Defective sequences with hypermutations were the major source of DRMs, with a median DRM rate of 1.0 compared with defective sequences without hypermutations (0.042, P&#x200a;<&#x200a;0.001). Certain Apolipoprotein B Editing Complex 3-related DRMs including reverse transcriptase gene mutations M184I, E138K, M230I, G190E and protease gene mutations M46I, D30N were enriched in hypermutated sequences but not in intact sequences or plasma sequences. All the hypermutated sequences had premature stop codons due to Apolipoprotein B Editing Complex 3. CONCLUSION: Proviral sequencing may overestimate DRMs as a result of hypermutations. Removing hypermutated sequences is essential in the interpretation of proviral drug resistance testing.

Anti-HIV Agents

Genomic detection of Panton-Valentine Leucocidins encoding genes, virulence factors and distribution of antiseptic resistance determinants among Methicillin-resistant S. aureus isolates from patients attending regional referral hospitals in Tanzania.

BACKGROUND: Methicillin-resistant Staphylococcus aureus (MRSA) is a formidable public scourge causing worldwide mild to severe life-threatening infections. The ability of this strain to swiftly spread, evolve, and acquire resistance genes and virulence factors such as pvl genes has further rendered this strain difficult to treat. Of concern, is a recently recognized ability to resist antiseptic/disinfectant agents used as an essential part of treatment and infection control practices. This study aimed at detecting the presence of pvl genes and determining the distribution of antiseptic resistance genes in Methicillin-resistant Staphylococcus aureus isolates through whole genome sequencing technology. MATERIALS AND METHODS: A descriptive cross-sectional study was conducted across six regional referral hospitals-Dodoma, Songea, Kitete-Kigoma, Morogoro, and Tabora on the mainland, and Mnazi Mmoja from Zanzibar islands counterparts using the archived isolates of Staphylococcus aureus bacteria. The isolates were collected from Inpatients and Outpatients who attended these hospitals from January 2020 to Dec 2021. Bacterial analysis was carried out using classical microbiological techniques and whole genome sequencing (WGS) using the Illumina Nextseq 550 sequencer platform. Several bioinformatic tools were used, KmerFinder 3.2 was used for species identification, MLST 2.0 tool was used for Multilocus Sequence Typing and SCCmecFinder 1.2 was used for SCCmec typing. Virulence genes were detected using virulenceFinder 2.0, while resistance genes were detected by ResFinder 4.1, and phylogenetic relatedness was determined by CSI Phylogeny 1.4 tools. RESULTS: Out of the 80 MRSA isolates analyzed, 11 (14%) were found to harbor LukS-PV and LukF-PV, pvl-encoding genes in their genome; therefore pvl-positive MRSA. The majority (82%) of the MRSA isolates bearing pvl genes were also found to exhibit the antiseptic/disinfectant genes in their genome. Moreover, all (80) sequenced MRSA isolates were found to harbor SCCmec type IV subtype 2B&5. The isolates exhibited 4 different sequence types, ST8, ST88, ST789 and ST121. Notably, the predominant sequence type among the isolates was ST8 72 (90%). CONCLUSION: The notably high rate of antiseptic resistance particularly in the Methicillin-resistant S. aureus strains poses a significant challenge to infection control measures. The fact that some of these virulent strains harbor the LukS-PV and LukF-PV, the pvl encoding genes, highlight the importance of developing effective interventions to combat the spreading of these pathogenic bacterial strains. Certainly, strengthening antimicrobial resistance surveillance and stewardship will ultimately reduce the selection pressure, improve the patient's treatment outcome and public health in Tanzania.

Methicillin-Resistant Staphylococcus aureus

ECHO: a nanopore sequencing-based workflow for (epi)genetic profiling of the human repeatome.

SUMMARY: The human genome is dominated by repetitive DNA, whose genetic and epigenetic variation plays a key role in gene regulation, genome stability, and disease. Recent advances in long-read sequencing now enable large-scale, haplotype-resolved, and DNA methylation-informative analysis of the human genome, including on previously inaccessible complex and repetitive regions. However, the comprehensive, simultaneous characterisation of the "human repeatome" remains challenging, largely due to the lack of comprehensive tools integrated in a single pipeline that can capture the full spectrum of variation across diverse types of DNA repeats. Here, we present ECHO, a user-friendly, Snakemake-based pipeline for the "(Epi)genomic Characterisation of Human Repetitive Elements using Oxford Nanopore Sequencing." ECHO provides a reproducible and scalable framework for end-to-end analysis of whole-genome nanopore sequencing data, enabling integrative but also tailored (epi)genetic analyses of the human repeatome. AVAILABILITY AND IMPLEMENTATION: ECHO is freely available at Github: https://github.com/leenput/ECHO-pipeline, with the archived version at Zenodo: https://zenodo.org/records/19068468.

Humans

Precision ID mtDNA Whole Genome Panel and sequencing of telogen hairs - perspectives for validation and implementation in casework.

Shed hair is a commonly encountered type of forensic evidence. Shed telogen hairs generally contain insufficient or highly degraded nuclear DNA for STR profiling; however, mtDNA analysis of telogen hair and hair shafts remains possible. We validated whole mitochondrial genome (mtGenome) sequencing using the Precision ID mtDNA Whole Genome Panel (Thermo Fisher Scientific) and subsequently implemented the panel for the analysis of telogen hair, buccal, and casework samples. We analysed 90 diluted DNA samples containing 3-3,600 mtDNA copies, shed telogen hairs and their corresponding mtDNA from buccal swabs from 91 individuals, and 11 archived DNA extracts from hair samples in criminal cases. Complete mtGenome sequences were consistently recovered in 99% of samples across DNA dilution series at DNA input levels as low as 47 mtDNA copies, demonstrating the assay's robustness under low-template conditions. We obtained complete and reproducible mtGenome sequences with &#x2265;&#x2009;327 mtDNA copies/&#xb5;L from telogen hair samples. After applying ISFG recommendations and excluding low-confidence discrepancies associated with high-strand bias, heteroplasmic variants and sequencing artifacts, mtGenome sequence concordance increased from 93.4% to 100%. None of the 16 negative controls produced complete mtDNA sequences. Six negative controls showed low-level mtDNA signal (2-8 variants), consisting predominantly of common polymorphisms. These samples did not yield complete mtGenome sequences and showed no correspondence to any of the analysed samples. Finally, archived telogen hair samples from criminal cases presented complete mtGenome sequences with an average read depth of 1,037x.Our findings highlight the reliability of mtDNA analysis of telogen hairs using the Precision ID mtDNA Whole Genome Panel for implementation in forensic casework.

Forensic casework

Genomic mapping of diabetic kidney disease biomarkers and identification of potential inhibitors through virtual screening.

BACKGROUND: Diabetic kidney disease (DKD) is a common and serious complication of diabetes mellitus, marked by a multifactorial pathogenesis and the absence of sensitive diagnostic biomarkers. Identifying novel molecular targets and therapeutic options is essential to improve early diagnosis and treatment outcomes. METHODS: To uncover potential biomarkers and therapeutic candidates, we performed an integrated genomic analysis using microarray and RNA-seq datasets from the Gene Expression Omnibus (GEO) and Sequence Read Archive (SRA) databases. Differentially expressed genes (DEGs) were identified and subjected to protein-protein interaction (PPI) network analysis. Key genes were further explored through virtual screening of an FDA-approved compound library using molecular docking techniques. Drug-likeness was assessed via Lipinski's rule of five. RESULTS: A total of 40 DEGs were identified, among which ISCU (downregulated; involved in iron-sulfur cluster biogenesis) and AP1S2 (upregulated; associated with vesicular trafficking) emerged as potential biomarkers. PPI analysis revealed their involvement in critical DKD-related pathways, such as extracellular matrix remodeling and oxidative stress. Virtual screening identified six FDA-approved compounds with high binding affinity (&#x2264;-7.96 kcal/mol) to ISCU, notably ZINC000001576020, all of which complied with Lipinski's rule. CONCLUSIONS: This in-silico study nominates ISCU and AP1S2 as candidate diagnostic biomarkers for DKD and identifies computationally prioritized inhibitors targeting ISCU. These findings require experimental validation but provide a molecular framework for precision diagnosis and therapeutic development. These findings offer new molecular insights that could inform precision diagnosis and personalized treatment strategies for diabetic kidney disease.

Diabetic Nephropathies

Gencube: centralized retrieval and integration of multi-omics resources from leading databases.

MOTIVATION: The volume of multi-omics data for diverse species is growing at an unprecedented rate, with new genome assemblies, related annotations, and high-throughput sequencing resources being submitted daily to various genomic data repositories. In response to this data influx, both existing and new databases are establishing optimized hierarchical structures to manage the vast amount of information. However, the lack of accessible command-line tools, combined with the functional limitations and unintuitive design of existing options, presents significant challenges for researchers. This gap underscores a critical need for a tool that enables streamlined retrieval and integration of omics data across these diverse repositories. RESULTS: We have developed Gencube, a command-line tool that enables centralized retrieval and integration of a comprehensive set of six different data types-genome assemblies, gene sets, annotations, sequences, comparative genomic data, and NGS-based omics resources-from various leading databases. AVAILABILITY AND IMPLEMENTATION: Gencube is a free and open-source tool, with its code available on GitHub: https://github.com/snu-cdrc/gencube and also archived on Zenodo: https://doi.org/10.5281/zenodo.14607649.

Databases, Genetic

A-liner: linear alignment visualizer for genome comparisons.

SUMMARY: A-liner is a flexible command-line tool for linear visualization of genome-scale sequence alignments, supporting outputs from multiple aligners and integrated visualization of annotations, highlights, quantitative tracks, and coordinate scales. It is applicable to a wide range of organisms, from bacteria to large eukaryotic genomes, and facilitates efficient generation of publication-ready comparative genome visualizations. AVAILABILITY AND IMPLEMENTATION: The source code and example output files for a-liner are available in the GitHub repository: https://github.com/mokuno3430/a-liner. A-liner v1.1.0 has been archived on Zenodo at https://doi.org/10.5281/zenodo.19702001.

Software

Sharkmer: repurposing PCR primers for targeted genome assembly using in silico PCR.

SUMMARY: We introduce an in silico PCR (sPCR) method for the assembly of specific genomic regions spanned by PCR primers using raw sequence reads. This allows a user to quickly isolate the exact regions that are abundant in public archives of gene sequences, leveraging the decades of work that have gone into optimizing primer sequences for benchtop PCR. We implement sPCR in sharkmer as a targeted de Bruijn graph assembler seeded with the forward primer sequence and terminated with the reverse primer sequence. This is useful for a variety of routine tasks, including validating the species identity of a dataset, identifying contaminants, and quickly building phylogenies from raw sequence data. AVAILABILITY AND IMPLEMENTATION: sharkmer is written in Rust. Code, instructions for installation and use, tests, and other resources are available in the GitHub repository at https://github.com/caseywdunn/sharkmer and at Zenodo with DOI 10.5281/zenodo.19020708. It can also be installed via bioconda.

Software

Generating multiple alignments on a pangenomic scale.

MOTIVATION: Since novel long read sequencing technologies allow for de novo assembly of many individuals of a species, high-quality assemblies are becoming widely available. For example, the recently published draft human pangenome reference was based on assemblies composed of contigs. There is an urgent need for a software-tool that is able to generate a multiple alignment of genomes of the same species because current multiple sequence alignment programs cannot deal with such a volume of data. RESULTS: We show that the combination of a well-known anchor-based method with the technique of prefix-free parsing yields an approach that is able to generate multiple alignments on a pangenomic scale, provided that large-scale structural variants are rare. Furthermore, experiments with real world data show that our software tool PANgenomic Anchor-based Multiple Alignment significantly outperforms current state-of-the art programs. AVAILABILITY AND IMPLEMENTATION: Source code is available at: https://gitlab.com/qwerzuiop/panama, archived at swh:1:dir:e90c9f664995acca9063245cabdd97549cf39694.

Software

Clinically actionable genomic alterations in breast cancer brain metastases.

BACKGROUND: Breast cancer brain metastases (BCBMs) represent a critical unmet clinical need in metastatic breast cancer (MBC) and the identification of novel therapeutic targets is urgently needed in this context. In this study, we describe clinically actionable targets in BCBMs using comprehensive genomic profiling. PATIENTS AND METHODS: Genomic DNA was extracted from formalin-fixed paraffin-embedded archival BCBM samples and analyzed using the commercially available Agilent SureSelect V6 whole exome sequencing (WES) kit and an Illumina NovaSeq 6000 platform. Pathogenic alterations were classified as actionable alterations (AAs) if they met the updated MBC or tumor-agnostic ESMO Scale for Clinical Actionability of Molecular Targets (ESCAT) I or II criteria of the ESCAT scale. RESULTS: WES data from 56 BCBM samples were available [33.9% hormone receptor (HR)-negative/human epidermal growth factor receptor (HER)2-negative; 25.0% HR-positive/HER2-negative; and 38% HER2-positive]. ESCAT I/II AAs were detected in 76.8% (n = 43) of all BCBMs and the most frequently detected AAs were in genes involved in the homologous recombination repair pathway (BRCA1/BRCA2/PALB2; 53.6% overall). Biallelic inactivation of BRCA1, BRCA2, or PALB2 was observed in 19.6% of samples, with higher rates in HER2-negative BCBMs (26% in HR-negative /HER2-negative and 21% in HR-positive/HER2-negative). ESCAT I/II PIK3CA/AKT1/PTEN pathway alterations were present in 48.2% of samples and, in particular, in 50% of HR-positive/HER2-negative BCBMs. No ESR1 mutation was detected in HR-positive/HER2-negative BCBMs. The prognostic impact of previously described AAs was evaluated overall and according to breast cancer subtype. Twenty-three BCBMs (41%) were classified as HER2-positive; among these, 3 (13%) presented a hotspot PIK3CA mutation and 7 (30%) presented a PTEN deletion. Among patients with HER2-positive BCBMs, the identification of a hotspot PIK3CA mutation was significantly associated with worse prognosis. CONCLUSIONS: ESCAT I/II actionable genomic alterations are frequent in BCBMs, highlighting the potential for genomically targeted treatments in this setting.

ESCAT