PubMed HealthSearch

SEARCH · PubMed Health

Results for “Sequence Count Data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

QCatch: a framework for quality control assessment and analysis of single-cell sequencing data.

MOTIVATION: Single-cell sequencing data analysis requires robust quality control (QC) to mitigate technical artifacts and ensure reliable downstream results. While tools like alevin-fry and simpleaf (and augmented execution context for the alevin-fry), offer flexibility and computational efficiency to process single-cell data, this ecosystem will further benefit from a standardized QC reporting tailored for its outputs. RESULTS: We introduce QCatch, a Python-based command-line tool that generates comprehensive and interactive HTML QC reports designed specifically for single-cell quantification results. Taking the output directory of alevin-fry or simpleaf as the input, QCatch is able to perform essential processing steps, like cell calling, and generate detailed QC reports that contain informative visualizations and statistics, including unique molecular identifier (UMI) count distributions, sequencing saturation estimates, and splicing status information, for QC assurance. Built for seamless integration into downstream analysis workflows, QCatch exports the processed results in a richly-annotated H5AD format file, a widely used data format common among many downstream single-cell data analysis tools. AVAILABILITY AND IMPLEMENTATION: The source code and documentation of QCatch are available on GitHub at https://github.com/COMBINE-lab/QCatch. QCatch can be installed via both Bioconda and PyPI.

Single-Cell Analysis

In vitro and in vivo studies on the impact of the familial adenomatous polyposis heterogeneous mutation MUC20-S671C on colorectal carcinogenesis and progression.

BACKGROUND: Familial adenomatous polyposis (FAP) is a hereditary colorectal cancer (CRC). We performed genetic testing on nine FAP patients and identified a recurrent mutation at the 671st site of the MUC20 gene-MUC20-S671C. This mutation has a detection frequency of zero in the 1000 Genomes Project database. Previous studies have demonstrated that MUC20 can promote CRC progression through epithelial-mesenchymal transition (EMT). We conducted a series of experiments to analyze the impact of this mutation on CRC cells, aiming to infer its potential role and significance in CRC patients. METHODS: We introduced the MUC20-S671C mutation into the CRC SW480 cell line using the CRISPR-Cas9 technique and established a stable cell line carrying this mutation. We then conducted various experiments to assess the effects of this mutation. The Transwell assay was used to evaluate cell invasion and migration. We also examined cell proliferation, cell cycle progression, and apoptosis rate. Furthermore, we tested the tumorigenic ability of these cells in NOD-scid IL2Rγ[null] (NSG) mice. Additionally, transcriptome sequencing was performed on both cell lines and mouse tumor tissues to obtain molecular regulatory network data, and key molecules were further validated. RESULTS: The results of Cell Counting Kit-8 (CCK-8), 5-ethynyl-2'-deoxyuridine (EdU), and colony formation assays indicated that the proliferation ability of mutant cells was significantly reduced. The Transwell assay demonstrated a marked decline in the invasion and migration capabilities of mutant cells. Flow cytometry analysis revealed that the mutation increased the apoptosis rate of CRC cells and might have caused S-phase arrest. The tumor formation assay in nude mice showed that the tumorigenic ability of mutant cells was weakened. Transcriptome sequencing of both the cells and tumor tissues suggested that the mutation altered the expression of apoptosis- and cell cycle-related molecules and also affected EMT. Further experiments confirmed that key molecules involved in the EMT process, such as E-cadherin, were upregulated, while Vimentin, MMP9, and MMP14 were significantly downregulated, indicating that the mutation weakened the EMT capability of CRC cells. CONCLUSIONS: We have identified a novel mutation, MUC20-S671C, in patients with FAP. Our study demonstrates that this mutation exerts its tumor-suppressive effect by reversing the EMT process.

MUC20-S671C

Non-structural maintenance of chromosome condensin I complex subunit H knockdown suppresses malignant progression of esophageal squamous cell carcinoma via the Wnt/β-catenin signaling pathway.

BACKGROUND: Esophageal squamous cell carcinoma (ESCC) remains a major cause of cancer-related mortality, and effective therapeutic targets are still limited. Non-structural maintenance of chromosome condensin I complex subunit H (NCAPH) has been implicated in tumorigenesis; however, its clinical relevance, functional roles, and underlying mechanisms in ESCC are not fully defined. We aimed to characterize the expression pattern, prognostic value, biological functions, and mechanistic basis of NCAPH in ESCC. METHODS: Public datasets from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) were analyzed to evaluate NCAPH expression and clinical associations. Single-cell RNA sequencing (scRNA-seq) data were used to map cell-type-specific distribution of NCAPH in tumor and adjacent tissues. NCAPH was silenced in KYSE150 and KYSE510 cells using lentiviral short hairpin RNAs (shRNAs), followed by Cell Counting Kit-8 (CCK-8), colony formation, wound-healing, and Transwell migration/invasion assays. A nude mouse xenograft model was established to assess the effect of NCAPH knockdown in vivo. RNA sequencing (RNA-seq), quantitative polymerase chain reaction (qPCR), western blotting, and enzyme-linked immunosorbent assay (ELISA) were performed to explore potential mechanisms. RESULTS: NCAPH was consistently upregulated in ESCC across multiple cohorts and was associated with unfavorable clinicopathological features and poorer survival. Functional assays demonstrated that NCAPH knockdown significantly inhibited ESCC cell proliferation, migration, invasion, and clonogenic growth. In vivo, NCAPH silencing suppressed xenograft tumor growth. Mechanistically, transcriptomic profiling and molecular validation indicated attenuation of Wnt/β-catenin signaling following NCAPH depletion, accompanied by reduced β-catenin and downstream targets. CONCLUSIONS: NCAPH promotes malignant progression of ESCC, at least in part through activation of the Wnt/β-catenin pathway, and may serve as a potential biomarker and therapeutic target.

Esophageal squamous cell carcinoma (ESCC)

Dual β-lactam therapy against high-risk Pseudomonas aeruginosa isolates: a dynamic in-vitro infection model study integrating population genomics with quantitative systems pharmacology modelling and simulations.

BACKGROUND: Pseudomonas aeruginosa has an extraordinary capacity for resistance emergence during treatment, even with newer antipseudomonals. There is a gap in understanding how resistance mechanisms affect the time-course of bacterial response to these newer agents. Traditional approaches for predicting pathogen response to an antibiotic do not apply to combination therapy. We aimed to develop a modelling framework to predict treatment response based on resistome information, using isolates of the worldwide-disseminated high-risk clone sequence type (ST) 235 and β-lactam antibiotics as the example. METHODS: In this hollow-fibre in-vitro infection study, we used three extensively drug-resistant ST235 clinical isolates from the national collection of the Clinical Microbiology Department of the Hospital Son Espases (Palma de Mallorca, Spain) that were hospital-acquired, were isolated following routine microbiological procedures from different patients between 2017 and 2022, were susceptible to ceftolozane-tazobactam, and had different levels of meropenem resistance. The selected isolates (ST235-05, ST235-09, and ST235-10) showed classical β-lactam resistance mechanisms pre-treatment. The isolates were investigated in 240-h dynamic hollow-fibre in-vitro infection models (HFIMs). The studies exposed the isolates to pharmacokinetic profiles of ceftolozane-tazobactam (simulating 1 g of ceftolozane and 0·5 g of tazobactam as a 3-h infusion every 8 h) and meropenem (simulating 6 g per day continuous infusion) as observed in hospitalised patients, as monotherapy and in combination. Treatment response was assessed through the quantification of the time-courses of viable total and resistant bacteria. Whole-genome sequencing identified the mechanisms of emerging resistance. A quantitative systems pharmacology (QSP) approach was used to model total and resistant bacterial counts and corresponding pharmacokinetic data from the HFIM. Monte Carlo simulations were used to predict treatment responses in 1000 virtual infected patients treated with ceftolozane-tazobactam and meropenem as monotherapies or in combination over 10 days. FINDINGS: In the HFIMs, each antibiotic alone amplified resistance by approximately 48 h for all isolates; that is, monotherapies resulted in a higher concentration of resistant bacteria compared with the control treatment at the respective time, except ceftolozane-tazobactam against ST235-10. Combination of ceftolozane-tazobactam and meropenem was synergistic (bacterial counts ≥2 log10 colony forming units [CFU] per mL lower than the best performing monotherapy and initial inoculum) against all isolates and suppressed resistance. Against ST235-10, ceftolozane-tazobactam monotherapy reduced counts to less than 1 log10 CFU per mL from 192 h onwards, whereas the combination reached less than 1 log10 CFU per mL by 24 h. Across strains, population genomics confirmed monotherapy failures were associated with emerging resistance mechanisms (ceftolozane-tazobactam: ampC Ω-loop mutations; meropenem: ftsl mutation). The developed QSP model incorporated baseline resistance mechanisms and those emerging in resistant mutant subpopulations. The model explained and predicted the monotherapy failures involving amplification of these subpopulations, and synergistic killing and resistance suppression by the combination. Simulations using the model predicted bacterial regrowth above the initial inoculum for more than 90% of patients after 0 to approximately 3 days for meropenem monotherapy across all strains and for ceftolozane-tazobactam monotherapy against ST235-05 and ST235-09. For ceftolozane-tazobactam monotherapy against ST235-10, regrowth was predicted for approximately 30% of patients. In contrast, the simulations predicted sustained bacterial killing of at least 2 log10 CFU per mL compared with the initial inoculum by the combination for more than 89% of patients across all strains. INTERPRETATION: To our knowledge, this model is the first to characterise and predict the time-course of responses of clinical isolates to antibiotics only by the resistance mechanisms present and their complex interplay, representing a step towards pathogen-specific, personalised medicine. FUNDING: Australian National Health and Medical Research Council.

Pseudomonas aeruginosa

An end-to-end computational framework for "Record-seq" transcriptional recording data.

MOTIVATION: Record-seq captures cumulative transcriptional activity over time in engineered Escherichia coli by integrating cellular RNA-derived spacer sequences into clustered regularly interspaced short palindromic repeats (CRISPR) arrays, which are read out by sequencing. Unlike the approximately uniform transcript sampling of RNA-seq, Record-seq records biological signal as spacers sampled by the CRISPR spacer acquisition machinery. Consequently, standard RNA-seq analysis strategies are not directly applicable, limiting sensitivity and interpretability. Our previous pipeline addressed these challenges only partially, retained inherited RNA-seq assumptions, and had limited algorithmic efficiency. RESULTS: Here, we present an end-to-end computational framework for Record-seq data. To address the primary computational bottleneck of spacer sequence extraction, we implemented a wavefront alignment approach for efficient quasi-local pattern matching, achieving an approximately 30-fold speedup. We introduce transcription unit-based feature counting as an alternative to gene-body quantification to better represent prokaryotic transcription and increase statistical power by capturing signal from untranslated regions, which are spacer acquisition hotspots. For downstream analyses, we incorporate multiple normalization strategies and a nonparametric differential expression testing framework designed for sparse datasets. Further, we analyze spacer acquisition patterns and train sequence-based neural models that predict acquisition propensity from genomic sequence and annotations, providing a framework for assessing whether acquisition rules generalize as Record-seq is extended to new microbial hosts. AVAILABILITY AND IMPLEMENTATION: The primary analysis workflow, the recoRdseq package, acquisition modeling repository, and relevant data are all linked at https://github.com/plattlab/Record-seq-Framework. Acquisition models and training data are on Zenodo at https://doi.org/10.5281/zenodo.18891434.

Escherichia coli

Modtector: ultra-fast modification signal mining on mapped sequencing reads.

SUMMARY: Existing tools for RNA epitranscriptomic modification and structural signal analysis are often fragmented, inefficiency, and limited to single signal types. We developed Modtector, an unified tool for extracting mutation and reverse-transcription stop signals from aligned sequencing reads. By using a "count-then-correct" strategy, Modtector reduces computational complexity and enables efficient dual-signal analysis. It achieves multi-fold speedups on large-genome and high-coverage datasets, including completing HEK293 22G data analysis in 5 minutes, and show strong scalability on single-cell datasets with speedups exceeding 50-fold. AVAILABILITY: The source code is available at GitHub (https://github.com/TongZhou2017/modtector) and Crates.io (https://crates.io/crates/modtector). The archived source-code snapshot used in this study is available at Zenodo (DOI: 10.5281/zenodo.20967747), corresponding to GitHub commit 7c60e9d. Workflow examples, datasets, and analysis scripts are available at Zenodo (DOI: 10.5281/zenodo.17316476 and 10.5281/zenodo.18523297).

Humans

Rare damaging CCR2 variants are associated with lower lifetime cardiovascular risk.

BACKGROUND: Previous work has shown a role of CCL2, a key chemokine governing monocyte trafficking, in atherosclerosis. However, it remains unknown whether targeting CCR2, the cognate receptor of CCL2, provides protection against human atherosclerotic cardiovascular disease. METHODS: Computationally predicted damaging or loss-of-function (REVEL > 0.5) variants within CCR2 were detected in whole-exome-sequencing data from 454,775 UK Biobank participants and tested for association with cardiovascular endpoints in gene-burden tests. Given the key role of CCR2 in monocyte mobilization, variants associated with lower monocyte count were prioritized for experimental validation. The response to CCL2 of human cells transfected with these variants was tested in migration and cAMP assays. Validated damaging variants were tested for association with cardiovascular endpoints, atherosclerosis burden, and vascular risk factors. Significant associations were replicated in six independent datasets (n = 1,062,595). RESULTS: Carriers of 45 predicted damaging or loss-of-function CCR2 variants (n = 787 individuals) were at lower risk of myocardial infarction and coronary artery disease. One of these variants (M249K, n = 585, 0.15% of European ancestry individuals) was associated with lower monocyte count and with both decreased downstream signaling and chemoattraction in response to CCL2. While M249K showed no association with conventional vascular risk factors, it was consistently associated with a lower risk of myocardial infarction (odds ratio [OR]: 0.66, 95% confidence interval [CI]: 0.54-0.81, p = 6.1 × 10-5) and coronary artery disease (OR: 0.74, 95%CI: 0.63-0.87, p = 2.9 × 10-4) in the UK Biobank and in six replication cohorts. In a phenome-wide association study, there was no evidence of a higher risk of infections among M249K carriers. CONCLUSIONS: Carriers of an experimentally confirmed damaging CCR2 variant are at a lower lifetime risk of myocardial infarction and coronary artery disease without carrying a higher risk of infections. Our findings provide genetic support for the translational potential of CCR2-targeting as an atheroprotective approach.

Humans

CNV-Finder: Streamlining Copy Number Variation Discovery.

Copy Number Variations (CNVs) play pivotal roles in the etiology of complex diseases and are variable across diverse populations. Understanding the association between CNVs and disease susceptibility is significant in disease genetics research and often requires analysis of large sample sizes. One of the most cost-effective and scalable methods for detecting CNVs is based on normalized signal intensity values, such as Log R Ratio (LRR) and B Allele Frequency (BAF), from Illumina genotyping arrays. In this study, we present CNV-Finder, a novel pipeline integrating deep learning techniques on array data, specifically a Long Short-Term Memory (LSTM) network, to expedite the large-scale identification of CNVs within predefined genomic regions. This facilitates efficient prioritization of samples for time-consuming or costly subsequent analyses such as Multiplex Ligation-dependent Probe Amplification (MLPA), short-read, and long-read whole genome sequencing. We incorporate four genes to establish our methods-Parkin (PRKN), Leucine Rich Repeat And Ig Domain Containing 2 (LINGO2), Microtubule Associated Protein Tau (MAPT), and alpha-Synuclein (SNCA)-which may be relevant to neurological diseases such as Alzheimer's disease (AD), Parkinson's disease (PD), Progressive Supranuclear Palsy (PSP), or related disorders such as essential tremor (ET). By training our models on expert-annotated samples and validating them across diverse cohorts, including those from the Global Parkinson's Genetics Program (GP2) and additional dementia-specific databases, we demonstrate the efficacy of CNV-Finder in accurately detecting deletions and duplications. Our pipeline outputs app-compatible files for visualization within CNV-Finder's interactive web application. This interface enables researchers to review predictions and filter displayed samples by model prediction values, LRR range, and variant count in order to explore or confirm results. Our pipeline integrates this human feedback to enhance model performance and reduce false positive rates. Through a series of comprehensive analyses and validations using visual inspection, MLPA, short-read, and long-read sequencing data, we demonstrate the robustness and adaptability of CNV-Finder in identifying CNVs with regions of varied size, probe density, and noise. Our findings highlight the significance of contextual understanding and human expertise in enhancing the precision of CNV identification, particularly in complex genomic regions like 17q21.31. The CNV-Finder pipeline is a scalable, publicly available resource for the scientific community, available on GitHub (https://github.com/GP2code/CNV-Finder; DOI 10.5281/zenodo.14182563). CNV-Finder not only expedites accurate candidate identification but also significantly reduces the manual workload for researchers, enabling future targeted validation and downstream analyses in regions or phenotypes of interest.

Copy Number Variation (CNV)

RP-REP Ribosomal Profiling Reports: an open-source cloud-enabled framework for reproducible ribosomal profiling data processing, analysis, and result reporting.

Ribosomal profiling is an emerging experimental technology to measure protein synthesis by sequencing short mRNA fragments undergoing translation in ribosomes. Applied on the genome wide scale, this is a powerful tool to profile global protein synthesis within cell populations of interest. Such information can be utilized for biomarker discovery and detection of treatment-responsive genes. However, analysis of ribosomal profiling data requires careful preprocessing to reduce the impact of artifacts and dedicated statistical methods for visualizing and modeling the high-dimensional discrete read count data. Here we present Ribosomal Profiling Reports (RP-REP), a new open-source cloud-enabled software that allows users to execute start-to-end gene-level ribosomal profiling and RNA-Seq analysis on a pre-configured Amazon Virtual Machine Image (AMI) hosted on AWS or on the user's own Ubuntu Linux server. The software works with FASTQ files stored locally, on AWS S3, or at the Sequence Read Archive (SRA). RP-REP automatically executes a series of customizable steps including filtering of contaminant RNA, enrichment of true ribosomal footprints, reference alignment and gene translation quantification, gene body coverage, CRAM compression, reference alignment QC, data normalization, multivariate data visualization, identification of differentially translated genes, and generation of heatmaps, co-translated gene clusters, enriched pathways, and other custom visualizations. RP-REP provides functionality to contrast RNA-SEQ and ribosomal profiling results, and calculates translational efficiency per gene. The software outputs a PDF report and publication-ready table and figure files. As a use case, we provide RP-REP results for a dengue virus study that tested cytosol and endoplasmic reticulum cellular fractions of human Huh7 cells pre-infection and at 6 h, 12 h, 24 h, and 40 h post-infection. Case study results, Ubuntu installation scripts, and the most recent RP-REP source code are accessible at GitHub. The cloud-ready AMI is available at AWS (AMI ID: RPREP RSEQREP (Ribosome Profiling and RNA-Seq Reports) v2.1 (ami-00b92f52d763145d3)).

AMI

An enhanced multisegment RT-PCR method for influenza A virus sequencing: Improved performance and reduced preparation time over traditional methods.

Influenza A viruses (IAVs) remain a major global health threat, affecting both human and animal populations. Whole-genome sequencing is essential for monitoring viral evolution, zoonotic transmission, and emerging variants. However, conventional RT-PCR methods often result in incomplete gene coverage, amplification biases, and reduced sequencing accuracy, particularly in clinical samples. We developed a robust In-house method for IAV full-genome sequencing using the Oxford Nanopore Technologies (ONT) long-read sequencing platform. This method integrates an in-house multisegment Reverse Transcription PCR (RT-PCR) method with a streamlined 2-pool primer design targeting all eight IAV gene segments. RNA extracted from clinical and stock virus samples was reverse-transcribed and amplified using Superscript IV-based chemistry, followed by magnetic bead purification to ensure high-quality amplicons. Sequencing libraries were prepared with the Native Barcoding Kit 24 (SQK-NBD114.24) and sequenced on R10.4.1 flow cells on the MinION MK1C device. Data analysis using the Iterative Refinement Meta-Assembler (IRMA) confirmed improved read depth, uniform coverage, and complete genome recovery. Compared to conventional methods, our In-House Multisegment 2-Pool (IH-MS2P) RT-PCR method generated higher numbers of matched read counts, minimized chimeric artifacts, and delivered superior genome coverage across human, swine, and avian isolates. This optimized RT-PCR method provides a high-performance, time-efficient, and portable solution for influenza genomics, demonstrating robust applicability even with clinical samples of low RNA yield.

Influenza A virus

Two Bacillus PGPB Strains in Wheat and Soybean: Wheat Growth Promotion Without Detectable Rhizosphere Microbiome Restructuring.

Plant growth-promoting bacteria (PGPB) are increasingly deployed as biofertilizers, yet the link between an inoculant's genomic potential and its realized effect on the plant is rarely assessed within an integrative framework that jointly captures the rhizosphere microbiome, plant phenotype, and strain genome. Two Bacillus strains-B. halotolerans 1453 and B. pumilus 630-were applied to wheat and soybean in a factorial pot experiment (2 strains &#xd7; 2 application methods &#xd7; 3 frequencies + control, 3-4 replicates). Rhizosphere samples (n = 67 after filtering) were profiled by 16S rRNA sequencing with PICRUSt2 functional prediction and compositional validation (Aitchison PERMANOVA, ALDEx2, ANCOM-BC2). The PGPB gene repertoire was characterized by genome mining (481 marker genes, 14 categories). Wheat phenotype (six traits) and soybean height were analyzed with models appropriate for count data (Negative Binomial and binomial GLMs) for treatment-vs.-control comparisons, and with factorial ANOVA for decomposition into main effects and interactions. Crop identity was the dominant factor shaping both microbiome structure and function (PERMANOVA R2 = 14.7% taxonomically and R2 = 7.8% functionally, both p < 0.001), with biologically meaningful taxonomic differences between wheat and soybean; strain, application count and method had no significant effect on community composition (R2 < 4% each), and co-occurrence networks showed no reliable differences between crops once read depth and sample size were controlled for. Despite this neutrality at the microbiome level, inoculation significantly increased wheat spike count (NB-GLM, all 12 treatments vs. control, padj 0.0002-0.031), ear weight, and stem count, with application count the strongest source of variability and a pronounced strain &#xd7; application count. Strain 1453 outperformed 630 in spike count (+23.1%, p = 0.012) and ear weight (+20.4%, p = 0.023); we hypothesize that this may be related to its more complete DNRA pathway (narGHI + nirB-nirD) and biocontrol genes (bacE, srfAA). Strain 630 produced a less pronounced effect than strain 1453 but was subject to smaller fluctuations across replicates (CV &#x2248; 16-21% vs. &#x2248;24-26% for 1453), which may reflect better resilience to environmental fluctuations, possibly due to its confirmed rsbV/rsbW stress-tolerance regulon. Rhizosphere microbiome composition differed clearly by crop (wheat vs. soybean) but showed no detectable response to strain, application method, or application count. Despite this lack of a microbiome signal, inoculation significantly increased wheat spike count and ear weight, with the magnitude and stability of this effect differing by strain. We hypothesize that this strain-dependent difference relates to underlying genomic differences-particularly in nitrogen metabolism (DNRA pathway) and stress-tolerance genes-though this link has not been tested directly and remains a hypothesis for future work.

Triticum

Telomere length and clonal hematopoiesis interact to influence outcomes in hematopoietic stem cell transplantation.

Clonal hematopoiesis (CH), the clonal expansion of a hematopoietic stem cell and its progeny driven by somatic mutations, has been associated with inferior survival outcomes among recipients of autologous stem cell transplants (ASCT). Leukocyte telomere length (LTL) has a complex but well-documented interaction with CH, but the impact of this interaction on stem cell transplantation has not been adequately examined. We measured LTL in graft cell DNA from 452 patients undergoing ASCT for myeloma, for whom targeted DNA sequencing for CH driver gene mutations was available. We interrogated clinical and longitudinal large-scale laboratory data for these patients to understand the impact of graft LTL on progression-free survival (PFS) and overall survival after transplantation, as well as blood count indices and their trajectories. In multivariate analyses, longer LTL was associated with increased PFS among patients without CH. However, this protective association was not seen in patients with CH. We also report that among patients with CH, longer LTL was associated with an increased red cell distribution width before myeloablative chemotherapy and after ASCT. Collectively, these data reveal hitherto undescribed interactions between LTL, CH, and ASCT outcomes.

Humans

De novo transcriptome assembly and gene expression analysis of Cnidium officinale under high-temperature conditions.

BACKGROUND: The medicinal plant Cnidium officinale (CO) is widespread in Northeast Asia and vulnerable to heat stress. The naturally occurring composition of pharmacological ingredients of CO results in overall physiological consequences; therefore, it is crucial to have a comprehensive understanding of metabolic response to ambient heat in terms of acclimation to estimate how much CO is exposed to threatening environmental conditions. RESULTS: Transcriptome analysis is critical for understanding the consequences of long-term physiological adaptation of CO to abiotic stress. However, transcriptome analysis on this species, particularly under prolonged stress conditions, has remained limited. We employed a temperature gradient tunnel (TGT) to subject CO to high-temperature exposure for four months, enabling us to observe the cumulative effects of heat and assess its acclimation mechanisms. In the absence of genome sequencing data, we performed de novo transcriptome assembly and compared DEGs from temperature treatment plots of a TGT and a growth chamber (GC). Since interpreting transcriptomic data can be complex, we employed a sequential analytical approach, including DEG clustering, GO enrichment, KEGG pathway mapping, miRNA-target gene analysis, and multiple rounds of RNA sequencing validation. DEGs were classified into two categories: genes exhibiting significant fold changes and genes showing significant count changes rather than fold changes. Then, we analyzed the functional roles&#xa0;of DEGs to determine which pathways respond to ambient and stressful high temperatures and validated the findings through cross-comparison with GC. Additionally, we conducted miRNA analysis to investigate post-transcriptional regulation under high temperatures. CO grown under higher ambient temperatures exhibited slight upregulation of pathways related to protein stability and turnover, ABA biosynthesis, and energy production, such as photosynthesis and oxidative phosphorylation. However, under extreme heat stress, most metabolic pathways were downregulated except for those involved in transcription, translation, oxidative phosphorylation and the biosynthesis of cutin, suberin, and wax. CONCLUSION: This study demonstrated that proper clustering of genes based on expression levels and fold changes in two different experimental conditions, along with pathway mapping, may provide a comprehensive understanding of CO's response to heat stress. These insights could contribute to future research on heat tolerance and crop improvement.

Gene Expression Profiling

Panaln: indexing pangenome for read alignment.

MOTIVATION: Pangenome indexing is a critical supporting technology in biological sequence analysis such as read alignment applications. The need to accurately identify billions of small sequencing fragments carrying sequencing errors and genomic variants drives the development of scalable and efficient pangenome indexing approach. RESULTS: We propose a new wavelet tree-based approach, called Panaln, for indexing pangenome and introduce a batch computation approach for fast count query over Panaln. We present a simple and effective seeding strategy and develop a pangenome program that uses the seed-and-extend paradigm for read alignment. Experimental results on simulated and real data demonstrate that Panaln uses significantly less space for the compared pangenome methods with generally higher accuracy. We provide a scalable index construction by representing pangenome with a linear model. Additionally, Panaln brings enhanced accuracy compared to the popular single reference methods. AVAILABILITY AND IMPLEMENTATION: Package: https://anaconda.org/bioconda/panaln and source code: https://github.com/Lilu-guo/Panaln.

Software

COSIGT: population-scalable genotyping of complex loci from low-coverage sequencing data using pangenome graphs.

Pangenome graphs capture extensive structural diversity, but resolving complex loci from shallow sequencing remains challenging, particularly when samples are of low quality such as in ancient DNA. We introduce COSIGT (COsine SImilarity-based GenoTyper), which assigns diploid genotypes by matching read-depth distributions to haplotype paths via cosine similarity. Because this metric evaluates relative coverage profiles rather than absolute read counts, COSIGT substantially outperforms existing likelihood-based tools at low coverage (1-2X). We demonstrate scalability to thousands of modern and ancient genomes, enabling robust, population-scale analyses of complex variation directly from low-coverage datasets.

Humans

Protein profiling and GC-MS product analysis provide insights into lignite solubilization and bioconversion by Lysinibacillus sphaericus strain SH19.

Lignite biosolubilization offers a mild route for valorizing low-rank coal, although the microbial processes that accompany solubilization remain incompletely defined. Here, an endogenous isolate designated Lysinibacillus sphaericus strain SH19 was evaluated using nitric-acid-pretreated Shengli lignite. Under the selected working conditions (4 M nitric-acid pretreatment, initial pH 8, 40&#xb0;C, and 16 days), the apparent solubilization rate reached 66.81%. Changes in A450, residual solid mass, culture pH, and extracellular protein concentration showed that chemical pretreatment and bacterial culture were both associated with the release of soluble lignite-derived material. SDS-PAGE and two-dimensional electrophoresis revealed treatment-associated differences in extracellular and intracellular protein patterns. LC-MS/MS analysis of excised protein spots yielded 85 candidate protein assignments; the revised supplementary table reports PEAKS scores, sequence coverage, peak area, and unique-peptide counts and highlights the limited support for several entries. GC-MS analysis produced 33 tentative library assignments in the solubilized fraction, but siloxane- and silyl-related signals were treated as possible analytical background, and no pathway was inferred from these assignments alone. Together, the data identify strain SH19 as a promising lignite-biosolubilizing isolate and provide candidate proteins and product signals for future validation. The proposed process model remains exploratory because direct enzyme assays, inhibitor experiments, carbon-balance measurements, transcriptomic or genetic validation, complete GC-MS blank subtraction, and authentic-standard confirmation were not available.

Bacillaceae

SISTEM: simulation of tumor evolution, metastasis, and DNA-seq data under genotype-driven selection.

SUMMARY: SISTEM is a software package and mathematical framework for simulating tumor evolution and cell migrations at single-cell resolution. Unlike existing frameworks which simulate cancer cell populations under the neutral coalescent or using simple birth-death models, SISTEM simulates tumor populations under somatic clonal selection using an agent-based framework. SISTEM can generate mutation profiles, read counts, and DNA sequencing reads along with ground truth cell lineages and migration graphs under a number of easily customizable mutation and selection models. For improved realism, SISTEM allows for cell fitness to be driven by genomic events of various scales including single nucleotide variants, segmental gains and losses, whole-chromosomal and chromosome-arm aberrations, and whole-genome duplications. SISTEM also includes numerous migration models to simulate metastatic cancers, facilitating the exploration and evaluation of diverse migration patterns. AVAILABILITY AND IMPLEMENTATION: SISTEM is written in Python and is freely available open-source under GNU GPLv3 from: https://github.com/samsonweiner/sistem.

Software

Functional characterization of the 9q34.13 locus identifies RAPGEF1 as modulating risk for melanoma and nevi via RAS activation.

Genome-wide association studies identified a melanoma- and nevus count-associated locus on chromosome band 9q34.13. Fine-mapping and melanocyte expression data collectively suggest two potential causal genes with opposite association with risk: higher levels of Rap guanine nucleotide exchange factor 1 (RAPGEF1) and lower levels of uridine-cytidine kinase 1 (UCK1). Colocalization analyses and conditional TWAS suggest multiple causal cis-regulatory sequence variants in partial linkage disequilibrium (LD) to each other. Melanocyte capture-HiC and CRISPR-inhibition demonstrated regulatory interactions between fine-mapped variants and the RAPGEF1 and UCK1 promoters. Focusing on RAPGEF1, we demonstrate RAPGEF1 expression promotes melanocyte growth and drives malignant transformation of human immortalized melanocytes. Following treatment with human EGF, RAPGEF1 overexpression activated both RAP1 and RAS. Further, we show RAPGEF1 expression is significantly enriched in melanomas lacking strongly activating RAS-MAPK mutations, suggesting that RAPGEF1 may promote oncogenic RAS-MAPK signaling in melanomas. Furthermore, in these tumors, we provide preliminary evidence to support the prognostic relevance of RAPGEF1 expression in patients lacking RAS or BRAF mutations. Together with other recent studies, these data suggest that germline variation influencing RAS activation may play a key role in nevus development and melanoma risk.

Journal Article