PubMed HealthSearch

SEARCH · PubMed Health

Results for “Sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Leveraging basecaller's move table to generate a lightweight k-mer model for nanopore sequencing analysis.

MOTIVATION: Nanopore sequencing by Oxford Nanopore Technologies (ONT) enables direct analysis of DNA and RNA by capturing raw electrical signals. Different nanopore chemistries have varied k-mer lengths, current levels, and standard deviations, which are stored in "k-mer models." In cases where official models are lacking or unsuitable for specific sequencing conditions, tailored k-mer models are crucial to ensure precise signal-to-sequence alignment, analysis and interpretation. The process of transforming raw signal data into nucleotide sequences, known as basecalling, is a fundamental step in nanopore sequencing. RESULTS: In this study, we leverage the move table produced by ONT's basecalling software to create a lightweight de novo k-mer model for RNA004 chemistry. We demonstrate the validity of our custom k-mer model by using it to guide signal-to-sequence alignment analysis, achieving high alignment rates (97.48%) compared to larger default models. Additionally, our 5-mer model exhibits similar performance as the default 9-mer models another analysis, such as detection of m6A RNA modifications. We provide our method, termed Poregen, as a generalizable approach for creation of custom, de novo k-mer models for nanopore signal data analysis. AVAILABILITY AND IMPLEMENTATION: Poregen is an open source package under an MIT license: https://github.com/hiruna72/poregen.

Nanopore Sequencing

Genetic predisposition to systemic inflammatory proteins is causally associated with inflammatory bowel disease: Insights from multi-omics association study and single-cell RNA-sequencing analysis.

Systemic inflammatory proteins have been reported to be related to inflammatory bowel disease (IBD) in previous observational research. However, their causal links remain obscure. Herein, we performed a Mendelian randomization (MR) analysis to analyze the causality between systemic inflammatory proteins and IBD. Genetic variants related to systemic inflammatory proteins were extracted from a meta-analysis of genome-wide association study (GWAS) data of 8293 European participants. Summary statistics of IBD diverse subtypes were obtained from the international IBD genetic consortium (IIBDGC). We conducted multi-omics method and MR study to detect the causal links through integrating GWAS and protein quantity trait loci (pQTL) data. Inverse variance weighted (IVW) approach was utilized as the dominated analysis method. Moreover, complementary approaches such as MR-Egger intercept test, Cochran Q test and leave-one-out analysis were utilized to validate pleiotropy and heterogeneity. Finally, single-cell RNA-sequencing analysis was performed to detect the expression of significant genes. For IBD, IVW estimates suggested that genetically predicted IL-10 and IL-13 were suggestively associated with an elevated risk of IBD (IL-10: OR: 1.12, 95% CI: 1.00-1.24, P = .04; IL-13: OR: 1.09, 95% CI: 1.01-1.18, P = .023), while CXCL10 was suggestively linked to a lower risk of IBD (CXCL10: OR: 0.90, 95% CI: 0.82-0.99, P = .037). For Crohn disease (CD), the IVW approach provided evidence to sustain that genetically determined IL-13 and CCL3 had a suggestive association with a higher risk of CD (IL-13: OR: 1.13, 95% CI: 1.02-1.26, P = .023; CCL3: OR: 1.22, 95% CI: 1.03-1.45, P = .018). Sensitivity analysis did not explore any heterogeneity and pleiotropy. Our findings supported the causal relationships between 4 specific inflammatory proteins (IL-10, IL-13, CXCL10, and CCL3) and the risk of IBD and CD, thereby providing promising biomarkers of various subtypes stratification and new insights for the prevention and therapeutic target of IBD.

Humans

A modular class-aware workflow for small RNA sequencing analysis using mouse sperm as a case study.

BACKGROUND: Small RNA sequencing analysis is challenging because RNA classes differ in biogenesis, sequence redundancy, genomic organization, and annotation reliability. Integrated workflows accommodating these constraints remain limited, particularly for fragment-level and cluster-level analysis. METHODS: We present a reproducible, containerized, class-aware workflow for small RNA sequencing analysis, using mouse sperm as a case study. The workflow combines standardized preprocessing with complementary annotation and quantification strategies for microRNAs (miRNAs), transfer RNA-derived small RNAs (tsRNAs), ribosomal RNA-derived small RNAs (rsRNAs), and PIWI-interacting RNA (piRNA)-enriched genomic clusters. Using sperm small RNA data from offspring of lipopolysaccharide (LPS)-exposed male mice, we compared integrated-reference mapping, multi-class annotation, fragment-level tsRNA profiling, and genome-based piRNA cluster analysis, with custom modules for locus-aware harmonization and condition-specific cluster analysis. RESULTS: Integrated-reference mapping aligned 88.17% of reads and retained 690 features after filtering. It identified 11 differentially expressed miRNAs between LPS and controls, while other classes showed limited signal. Fragment-level profiling improved tsRNA resolution. piRNA cluster analysis identified 958 control and 940 LPS clusters, with 18 control-specific and no LPS-specific clusters. CONCLUSION: This workflow supports transparent, reproducible, class-aware interpretation of small RNA sequencing data while emphasizing cautious interpretation of piRNA-enriched signals from total small RNA sequencing.

Small non-coding RNA analysis

The Hunt Lab Guide to De Novo Peptide Sequence Analysis by Tandem Mass Spectrometry.

Donald Hunt has made seminal contributions to the fields of proteomics, immunology, epigenetics, and glycobiology. The foundation of every important work to come out of the Hunt Laboratory is de novo peptide sequencing. For decades, he taught hundreds of students, postdocs, engineers, and scientists to directly interpret mass spectral data. To honor his legacy and ensure that the art of de novo sequencing is not lost, we have adapted his teaching materials into "The Hunt Lab Guide to De Novo Peptide Sequence Analysis by Tandem Mass Spectrometry". In addition to the de novo sequencing tutorials, we present two freely available software tools that facilitate manual interpretation of mass spectra and validation of search results. The first, "Hunt Lab Peptide Fragment Calculator", calculates precursor and fragment mass-to-charge ratios for any peptide. The second program, "Predator Protein Fragment Calculator", was inspired in part by the fragment calculator developed in the Hunt Lab. Its capabilities are enhanced to facilitate interpretation of mass spectral data derived from intact proteins. We hope that the combination of these educational tools will continue to benefit students and researchers by empowering them to interpret data on their own.

Tandem Mass Spectrometry

The impact of the COVID-19 pandemic on the incidence of invasive pneumococcal disease in the Czech Republic and whole genome sequencing analysis of Streptococcus pneumoniae serotypes 3 and 19A from 2018-2024.

AIM: To describe in detail changes in the incidence of invasive pneumococcal disease in the Czech Republic during and after the COVID-19 pandemic. Another objective is molecular analysis of S. pneumoniae isolates of serotypes 3 and 19A recovered in the Czech Republic between 2018 and 2024. MATERIAL AND METHODS: Data on the incidence of invasive pneumococcal disease and S. pneumoniae serotypes were obtained from the invasive pneumococcal disease surveillance program in the Czech Republic. S. pneumoniae isolates of serotypes 3 (63) and 19A (66) from 2018-2024 were subjected to whole genome sequencing (WGS) to characterize the GPSCs (Global Pneumococcal Sequence Clusters) and STs (sequence types) and place them in a global context. RESULTS: Results: During the COVID-19 pandemic, a significant decline was observed in the incidence of invasive pneumococcal disease in the Czech Republic. Following the pandemic, the incidence of invasive pneumococcal disease rose again to significantly higher levels than before the pandemic. Compared to the 2018–2019 period, the incidence of certain serotypes increased in 2023–2024, including vaccine serotypes 3, 4, 14, and 15B, while the incidence of serotypes 8, 12F, and 15A, among others, decreased. Whole genome sequencing analysis demonstrated the dominance of GPSC12 ST-180 among serotype 3 isolates throughout the study period. Among serotype 19A isolates, GPSC4 prevailed, particularly ST-416. CONCLUSIONS: The COVID-19 pandemic has demonstrated how rapidly the epidemiological situation of invasive pneumococcal disease can change and that continuous, systematic surveillance of invasive pneumococcal disease is necessary. The best prevention against invasive pneumococcal disease is vaccination, primarily with higher valency pneumococcal conjugate vaccines.

Czech Republic

SIMS: A deep-learning label transfer tool for single-cell RNA sequencing analysis.

Cell atlases serve as vital references for automating cell labeling in new samples, yet existing classification algorithms struggle with accuracy. Here we introduce SIMS (scalable, interpretable machine learning for single cell), a low-code data-efficient pipeline for single-cell RNA classification. We benchmark SIMS against datasets from different tissues and species. We demonstrate SIMS's efficacy in classifying cells in the brain, achieving high accuracy even with small training sets (<3,500 cells) and across different samples. SIMS accurately predicts neuronal subtypes in the developing brain, shedding light on genetic changes during neuronal differentiation and postmitotic fate refinement. Finally, we apply SIMS to single-cell RNA datasets of cortical organoids to predict cell identities and uncover genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Single-Cell Analysis

Whole genome sequence analysis of low-density lipoprotein cholesterol across 246&#xa0;K individuals.

BACKGROUND: Rare genetic variation provided by whole genome sequence datasets has been relatively less explored for its contributions to human traits. Meta-analysis of sequencing data offers advantages by integrating larger sample sizes from diverse cohorts, thereby increasing the likelihood of discovering novel insights into complex traits. Furthermore, emerging methods in genome-wide rare variant association testing further improve power and interpretability. RESULTS: Here, we conduct the largest meta-analysis of whole genome sequencing for low-density lipoprotein cholesterol (LDL-C), a therapeutic target for coronary artery disease, analyzing data from 246&#xa0;K participants and integrating 1.23B variants from the UK Biobank and the Trans-Omics for Precision Medicine (TOPMed) program. We identify numerous rare coding and non-coding gene associations related to LDL-C, with replication across 86&#xa0;K participants in All of Us. Our findings are based on single-variant analyses, rare coding and non-coding variant aggregation tests, and sliding window approaches. Through this comprehensive analysis, we identify 704 novel single-variant associations, 25 novel rare coding variant aggregates, 28 novel rare non-coding variant aggregates, and one novel sliding window aggregate. CONCLUSIONS: This study provides a meta-analysis framework for large-scale whole genome sequence association analyses from diverse population groups, yielding novel rare non-coding variant associations.

Humans

Dual RNA isolation from blood: an optimized protocol for host and bacterial RNA purification for dual RNA-sequencing analysis in whole blood sepsis samples.

Dual RNA-sequencing (dual RNA-seq) holds significant promise for deciphering bacterial virulence mechanisms during systemic infections. However, its application in sepsis research is hindered by technical challenges, including a low bacterial burden in blood and limited sample volumes and RNA yield from vulnerable populations, such as neonates. We developed an optimized protocol [dual RNA isolation from blood (DRIB)] for simultaneous stabilization, isolation and purification of high-quality host leukocyte and bacterial RNA from low-volume whole blood samples (0.5&#x2009;ml). This protocol is compatible with clinical sample collection workflows and high-throughput RNA sequencing. The feasibility of DRIB for dual RNA-seq was validated using a pilot cohort of clinical adult sepsis samples, enabling the investigation of host-bacterial gene expression during sepsis. The DRIB protocol yielded 2.10-6.91&#x2009;&#xb5;g of total RNA per clinical sample in our pilot cohort. Dual-species ribosomal RNA (rRNA) depletion and RNA-seq generated 16.6-24.8&#x2009;million filtered reads per sample, with 63&#xb1;7% of reads uniquely mapped to host or bacterial sequences. Host genes accounted for 51-68% (8.4-10.9&#x2009;million) reads, while 0.5-6.7% (79,496-789,808 reads) mapped to bacterial genomes. Bioinformatic analysis revealed that both shared and individual transcriptional patterns were identified in host and bacterial responses, including pathways related to immune metabolism and metal-ion binding. Our optimized DRIB protocol and RNA-seq pipeline effectively captured both host and bacterial RNA transcription in clinical sepsis samples. Expanding this approach to larger cohorts and varying disease timepoints will provide crucial new insights into host-bacterial gene co-expression dynamics in sepsis progression and outcomes.

Humans

CLN3 transcript complexity revealed by long-read RNA sequencing analysis.

BACKGROUND: Batten disease is a group of rare inherited neurodegenerative diseases. Juvenile CLN3 disease is the most prevalent type, and the most common pathogenic variant shared by most patients is the "1-kb" deletion which removes two internal coding exons (7 and 8) in CLN3. Previously, we identified two transcripts in patient fibroblasts homozygous for the 1-kb deletion: the 'major' and 'minor' transcripts. To understand the full variety of disease transcripts and their role in disease pathogenesis, it is necessary to first investigate CLN3 transcription in "healthy" samples without juvenile CLN3 disease. METHODS: We leveraged PacBio long-read RNA sequencing datasets from ENCODE to investigate the full range of CLN3 transcripts across various tissues and cell types in human control samples. Then we sought to validate their existence using data from different sources. RESULTS: We found that a readthrough gene affects the quantification and annotation of CLN3. After taking this into account, we detected over 100 novel CLN3 transcripts, with no dominantly expressed CLN3 transcript. The most abundant transcript has median usage of 42.9%. Surprisingly, the known disease-associated 'major' transcripts are detected. Together, they have median usage of 1.5% across 22 samples. Furthermore, we identified 48 CLN3 ORFs, of which 26 are novel. The predominant ORF that encodes the canonical CLN3 protein isoform has median usage of 66.7%, meaning around one-third of CLN3 transcripts encode protein isoforms with different stretches of amino acids. The same ORFs could be found with alternative UTRs. Moreover, we were able to validate the translational potential of certain transcripts using public mass spectrometry data. CONCLUSION: Overall, these findings provide valuable insights into the complexity of CLN3 transcription, highlighting the importance of studying both canonical and non-canonical CLN3 protein isoforms as well as the regulatory role of UTRs to fully comprehend the regulation and function(s) of CLN3. This knowledge is essential for investigating the impact of the 1-kb deletion and rare pathogenic variants on CLN3 transcription and disease pathogenesis.

Humans

Whole exome sequencing analysis of 167 men with primary infertility.

BACKGROUND: Spermatogenic failure is one of the leading causes of male infertility and its genetic etiology has not yet been fully understood. METHODS: The study screened a cohort of patients (n&#x2009;=&#x2009;167) with primary male infertility in contrast to 210 normally fertile men using whole exome sequencing (WES). The expression analysis of the candidate genes based on public single cell sequencing data was performed using the R language Seurat package. RESULTS: No pathogenic copy number variations (CNVs) related to male infertility were identified using the the GATK-gCNV tool. Accordingly, variants of 17 known causative (five X-linked and twelve autosomal) genes, including ACTRT1, ADAD2, AR, BCORL1, CFAP47, CFAP54, DNAH17, DNAH6, DNAH7, DNAH8, DNAH9, FSIP2, MSH4, SLC9C1, TDRD9, TTC21A, and WNK3, were identified in 23 patients. Variants of 12 candidate (seven X-linked and five autosomal) genes were identified, among which CHTF18, DDB1, DNAH12, FANCB, GALNT3, OPHN1, SCML2, UPF3A, and ZMYM3 had altered fertility and semen characteristics in previously described knockout mouse models, whereas MAGEC1,RBMXL3, and ZNF185 were recurrently detected in patients with male factor infertility. The human testis single cell-sequencing database reveals that CHTF18, DDB1 and MAGEC1 are preferentially expressed in spermatogonial stem cells. DNAH12 and GALNT3 are found primarily in spermatocytes and early spermatids. UPF3A is present at a high level throughout spermatogenesis except in elongating spermatids. The testicular expression profiles of these candidate genes underlie their potential roles in spermatogenesis and the pathogenesis of male infertility. CONCLUSION: WES is an effective tool in the genetic diagnosis of primary male infertility. Our findings provide useful information on precise treatment, genetic counseling, and birth defect prevention for male factor infertility.

Humans

Evaluating 12 automated, whole-genome sequencing analysis pipelines for Mycobacterium tuberculosis complex: a comparative study.

BACKGROUND: Reliance on complex, custom-built bioinformatics pipelines is a barrier to the implementation of whole-genome sequencing (WGS) of Mycobacterium tuberculosis in high-burden settings in some low-income and middle-income countries (LMICs). Automated analysis pipelines could address this inequity in access to WGS-based diagnostics and surveillance. This study aimed to systematically evaluate the performance and usability of publicly available WGS pipelines for M tuberculosis. METHODS: We identified automated M tuberculosis WGS analysis pipelines through searches of PubMed and GitHub from database inception up to Aug 31, 2024. Accuracy, cost, accessibility, and scalability were assessed for each pipeline. We evaluated the accuracy of genotypic drug susceptibility testing (gDST) using publicly available sequences with phenotypic susceptibility data for 12 antituberculosis drugs. We estimated pooled sensitivity and specificity for each pipeline, across all drugs, by conducting a bivariate meta-analysis, with random effects representing between-drug variability. Lineage classifications were compared, and a previously epidemiologically well-characterised dataset was used to compare measures of genomic relatedness. FINDINGS: Among 28 candidate pipelines, 16 were excluded as they were unmaintained and inexecutable. 12 pipelines (11 compatible with Illumina and four compatible with Nanopore), all free to use, were included for evaluation. Six pipelines processed and stored data remotely, but for five of these six, scalability was limited by the need to upload sequences through web portals. For local processing pipelines, scalability was dependent on substantial local computational resources, data storage capacity, and command-line interfaces that limited user-friendliness. Only one of six remote-processing pipelines removed human DNA sequences before server upload. gDST was similarly accurate across ten of 11 Illumina-compatible pipelines and three of four Nanopore-compatible pipelines. All pipelines classified the main lineages consistently, although there were differences at sublineage resolution. Outputs from three of four pipelines reporting genomic relatedness were compatible with commonly cited single nucleotide polymorphism difference thresholds. INTERPRETATION: Numerous automated analysis pipelines capable of enhancing equity in M tuberculosis WGS are available. Given the overall similarities between the pipelines evaluated in this study in terms of gDST performance, lineage classification, and genomic relatedness inference, non-functional attributes such as availability, accessibility, scalability, and privacy could represent the point of difference for prospective users in LMICs with a high burden of tuberculosis. FUNDING: The Rhodes Trust, Wellcome, Ellison Institute of Technology, and the UK National Institute for Health and Care Research Oxford Biomedical Research Centre.

Mycobacterium tuberculosis

Exploring biosynthetic potential of the endophytic Penicillium turbatum BLH34 using whole-genome sequence analysis and molecular networking.

An in-depth genomic and metabolomic investigation was conducted on the endophytic fungus Penicillium turbatum BLH34, isolated from Macleaya cordata. Hybrid sequencing (Illumina-Nanopore) generated a high-quality 27.9&#x2009;Mb genome (GC 48.6%) encoding 9798 proteins, with functional annotation linking 5350 genes to the NCBI non-redundant database and 3404 to KEGG pathways. AntiSMASH analysis uncovered 35 biosynthetic gene clusters (BGCs), 23 of which lacked homology to known pathways, highlighting BLH34's potential for novel metabolite discovery. Molecular networking (GNPS) and LC-MS/MS identified 19 specialised metabolites, including antimicrobial polyketides. Bioassays demonstrated potent inhibition against Staphylococcus aureus (36&#x2009;mm), Bacillus subtilis (28&#x2009;mm) and Escherichia coli (24&#x2009;mm), underscoring its pharmaceutical relevance.

Penicillium

Unraveling Neuronal Identities Using SIMS: A Deep Learning Label Transfer Tool for Single-Cell RNA Sequencing Analysis.

Large single-cell RNA datasets have contributed to unprecedented biological insight. Often, these take the form of cell atlases and serve as a reference for automating cell labeling of newly sequenced samples. Yet, classification algorithms have lacked the capacity to accurately annotate cells, particularly in complex datasets. Here we present SIMS (Scalable, Interpretable Machine Learning for Single-Cell), an end-to-end data-efficient machine learning pipeline for discrete classification of single-cell data that can be applied to new datasets with minimal coding. We benchmarked SIMS against common single-cell label transfer tools and demonstrated that it performs as well or better than state of the art algorithms. We then use SIMS to classify cells in one of the most complex tissues: the brain. We show that SIMS classifies cells of the adult cerebral cortex and hippocampus at a remarkably high accuracy. This accuracy is maintained in trans-sample label transfers of the adult human cerebral cortex. We then apply SIMS to classify cells in the developing brain and demonstrate a high level of accuracy at predicting neuronal subtypes, even in periods of fate refinement, shedding light on genetic changes affecting specific cell types across development. Finally, we apply SIMS to single cell datasets of cortical organoids to predict cell identities and unveil genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. When cell types are obscured by stress signals, label transfer from primary tissue improves the accuracy of cortical organoid annotations, serving as a reliable ground truth. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Brain organoids

The role of KIAA1467 in breast cancer: insights from pan-cancer and single-cell sequencing analysis.

BACKGROUND: Improving the response rate of single-agent immune checkpoint blockade (ICB) urgently requires the discovery of new therapeutic targets for combinatorial regimens. Analyses of tumor microenvironment (TME)-associated biomarkers have verified that KIAA1467 drives the formation of an immune-excluded, non-inflamed TME in breast cancer (BRCA). This study systematically explores the expression pattern, prognostic value, immune regulatory function, biological effects, and drug resistance relevance of FAM234B (also known as KIAA1467) in BRCA. METHODS: We performed pan-cancer survival analysis using The Cancer Genome Atlas (TCGA) datasets. Multi-omics bioinformatics analyses were conducted to evaluate KIAA1467 expression across malignancies. Single-cell RNA sequencing (scRNA-seq) data from GSE176078 was utilized to localize KIAA1467 expression at the cellular level. Immunohistochemistry and western blot assays validated KIAA1467 expression in BRCA clinical specimens. Correlation analyses were implemented to assess relationships between KIAA1467 expression, clinicopathological features, immune modulators, tumor-infiltrating immune cells, and p53 mutation status. Functional enrichment analysis uncovered relevant signaling pathways. Bioinformatic half maximal inhibitory concentration (IC50) prediction and in vitro cellular experiments were applied to evaluate associations between KIAA1467 and chemotherapeutic drug sensitivity. RESULTS: TCGA pan-cancer survival analysis demonstrated that elevated KIAA1467 expression significantly predicted shortened overall survival in BRCA and multiple other tumor types. KIAA1467 displayed distinct expression patterns across cancers, with prominent upregulation in BRCA. scRNA-seq confirmed enriched KIAA1467 expression within BRCA cells, and its upregulation in BRCA tissues was further verified by immunohistochemistry and western blot. High KIAA1467 expression was positively correlated with advanced tumor grade and lymphatic metastasis. KIAA1467 showed negative correlations with most immune modulators and core immune checkpoint molecules, as well as tumor-infiltrating immune cells in the TME, implying its potential function in tumor immune evasion. Low KIAA1467 expression was tightly linked to p53 mutations. Enrichment analysis indicated participation of KIAA1467 in epithelial-mesenchymal transition, apoptosis and cell cycle arrest. Furthermore, high KIAA1467 expression corresponded to higher estimated IC50 values of cisplatin, gefitinib, paclitaxel and gemcitabine, consistent with reduced chemosensitivity observed in vitro. CONCLUSIONS: This study reveals the multifaceted oncogenic role of KIAA1467 in BRCA. KIAA1467 participates in remodeling an immunosuppressive TME, correlates with malignant progression and chemoresistance, and may serve as a promising candidate target to optimize ICB-based combination therapy for BRCA. These findings offer new perspectives for the clinical treatment and comprehensive management of BRCA.

KIAA1467

Whole genome sequencing analysis of Candida glabrata isolates collected from patients with selected drug-resistant candidiasis hospitalized in Eastern Poland.

The epidemiology data for candidiasis indicate an increase in Candida glabrata infections. Moreover, several reports have shown an increasing number of drug-resistant cases of these infections. The source of drug resistance can often be traced to genetic mutations in genes related to a drug's mechanism of action. Therefore, we conducted whole genome sequencing of several drug-resistant isolates of Candida glabrata collected from patients hospitalized in Eastern Poland to assess whether mutations in selected genes correlated with susceptibility analysis results. The fungal species from patient samples were identified, and the isolated Candida glabrata were subjected to antifungal drug susceptibility testing. The results were interpreted according to the EUCAST and CLSI recommendations. Susceptibility to 5-flucytosine was assessed using the ATB FUNGUS kit. Libraries were prepared according to the NEXTERA XT DNA Library Prep and subsequently sequenced. The outcomes indicated common resistance to two of the three analyzed echinocandins, as well as two cases of simultaneous resistance to echinocandins and selected azole-based drugs. We detected several previously reported mutations in selected resistance-related genes, as well as five that are first described here: ERG5 (M267I), ERG6 (R57K), PDH1 (K438Q, V434I, F600V, V1192S), FCY1 (M129T), and FCY2 (I384F). Neither of the identified nonsynonymous mutations was correlated with the drug resistance demonstrated in the susceptibility testing. Furthermore, we can exclude the possibility of acquired drug resistance, thereby raising questions about the possibility of unknown mechanisms of resistance to azole-based and echinocandin drugs.

Candida glabrata

Whole-genome Sequence Analysis Revealed Novel Subjective Cognitive Decline-associated Genes in 10,763 Chinese.

Subjective cognitive decline (SCD) is widely regarded as a potential preclinical stage of Alzheimer's disease (AD), yet its genetic basis remains poorly understood. To address this gap, we investigated genetic biomarkers associated with SCD using whole-genome sequencing (WGS) in 10,763 Chinese participants from the Healthy Zhejiang One Million People Cohort (HOPE Cohort). The discovery stage included 9284 samples, with 1479 samples used for validation. Using a two-stage design, we systematically investigated both common and rare variants associated with SCD. In rare variant analyses, we identified and replicated an association between the upstream region of SEPHS2 and SCD. SEPHS2 is involved in selenophosphate synthesis, and a Mendelian randomization analysis reveals that its expression levels in both blood and brain cerebellum are associated with AD. Additionally, we identified CLVS2, which encodes a protein primarily expressed in neuronal cells, as a potential regulator for SCD based on missense rare variants. Multi-omics evidence suggests that both SEPHS2 and CLVS2 may play roles in neurodegenerative diseases. For common variants, we validated 8 known loci related to cognitive decline, 3 of which originated from the only existing SCD genetic study conducted under a migraine background. Overall, our WGS-based study fills the gap in SCD research by providing vital genetic evidence from an East Asian population and offers insights into the pathogenic mechanisms of SCD.

Aged

Transmission of extended spectrum &#x3b2;-lactamase-producing Escherichia coli and antimicrobial resistance gene flow across One Health compartments in eastern Africa: a whole-genome sequence analysis from a prospective cohort study.

BACKGROUND: The One Health paradigm considers interdependence of human, animal, and environmental health. However, there is little evidence from high-income countries to support the importance of a One Health approach to addressing spread of antimicrobial resistance (AMR). Given AMR is a global threat, understanding how the close interactions of humans with animals and the environment in low-income settings affect the spread of AMR is important. We aimed to investigate diversity and transmission of extended spectrum &#x3b2;-lactamase (ESBL)-producing Escherichia coli across household-linked One Health compartments using genomic data. METHODS: We sequenced whole genomes of ESBL-producing E coli isolates from humans, animals, and the environment from a prospective, longitudinal cohort study conducted in Malawi (April 29, 2019, to Dec 3, 2020) and Uganda (July 16, 2020, to Aug 6, 2021). In the cohort study, 259 households were enrolled at baseline in Malawi and 92 in Uganda from a mix of urban, peri-urban, and rural areas. Households were followed up at months 1, 3, and 6 in Malawi and at months 1, 2, and 4 in Uganda. Samples collected at each visit included human and animal stool, environmental samples from hand-contact areas, food, and water, and broader environmental samples such as river water. Samples were cultured in buffered peptone water and then ESBL chromogenic agar to isolate ESBL-producing E coli. ESBL-producing E coli isolates underwent whole-genome sequencing. We performed phylogenetic analyses, and in-silico multi-locus sequence typing, characterised AMR determinants and linked genotypes with sample location, ecological source, and other covariates. We performed fine-scale single nucleotide polymorphism (SNP) and network analysis to infer strain and plasmid transmission across ecological compartments. The primary outcome was colonisation with ESBL-producing E coli. Secondary outcomes were genomic clusters and ESBL genomic determinants within and between One Health compartments. FINDINGS: We found high diversity of ESBL-producing E coli, with 170 sequence types and 166 genomic clusters identified from 2344 genomes, including 1814 genomes from Malawi (907 human, 221 animal, and 686 environmental) and 530 genomes from Uganda (380 human, 147 animal, and three environmental). Sequence type (ST)131 dominated in Malawi (209 [11&#xb7;5%] of 1814 genomes), and ST10 dominated in Uganda (45 [8&#xb7;5%] of 530 genomes). Common ESBL genes blaCTX-M-15 (1604 [68&#xb7;4%] of 2344 genomes) and blaCTX-M-27 (336 [14&#xb7;3%] of 2344 genomes) were carried on a complex network of 55 and 30 different plasmids. This diversity of plasmids presented multiple pathways for dissemination and revealed high force of selection. Phylogenetic analyses revealed common intermixing of isolates between humans, animals, and the environment. SNP transmission analysis revealed ecologically overlapping clusters, suggesting ESBL-producing E coli co-circulation both within and between compartments with frequent spillover events. Applying a five-SNP threshold, we inferred 463 human-environment transmission events, 146 human-animal events, and 142 animal-environment events. INTERPRETATION: Our work suggests that a One Health approach is crucial to addressing AMR in eastern Africa. Improving water, sanitation, and hygiene systems will create a safer environment, reduce spillovers of AMR bacteria between compartments, and eventually reduce AMR reservoirs in the environment and in animals. FUNDING: Medical Research Council, National Institute for Health and Care Research, and Wellcome Trust.

Humans

Whole genome sequencing analysis and functional characterization of Lacticaseibacillus rhamnosus HP-B1083.

Lacticaseibacillus rhamnosus is an important strain for the biotransformation of natural products, and its crude extract exhibits biotransformation effect on glycosidic compounds such as baicalin. To further explore the potential of this strain, particularly given its previously demonstrated high-efficiency &#x3b2;-glucuronidase activity for baicalin conversion, whole-genome sequencing and functional annotation of Lacticaseibacillus rhamnosus HP-B1083 were performed in this study, and its acid tolerance, bile salt tolerance, short-term heat resistance and antibacterial activity were evaluated. The results showed that the strain possessed a circular chromosome with a full length of 3,090,505&#xa0;bp and a GC content of 46.69%. Gene annotation revealed that the genome contained 2941 coding sequences (CDS) and 112 non-coding RNA genes, including 60 tRNA genes, 1 tmRNA gene, 36 misc_RNA genes and 15 rRNA genes. The functional annotations further reveal that this genome is rich in genes related to carbohydrate metabolism, hydrolases, and transferases, which is highly consistent with its phenotypic characteristics in glycoside transformation and the synthesis of antibacterial substances. In addition, acid tolerance, bile salt tolerance and short-term heat resistance experiments verified that HP-B1083 had acid resistance, bile salt resistance and short-term heat resistance. Antibacterial activity tests confirmed that HP-B1083 produced inhibition zone diameters over 10&#xa0;mm against common foodborne pathogenic bacteria such as Escherichia coli and Bacillus cereus. Therefore, Lacticaseibacillus rhamnosus HP-B1083 has important application prospects in the development of functional foods, preparation of enzyme preparations and pharmaceutical industry.

Whole Genome Sequencing