PubMed HealthSearch

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A trainable language model with potential to modulate translation rates in non-model organisms by generating upstream untranslated region sequence libraries.

Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5'-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanistic assumptions, prior knowledge that may not be available in non-model contexts, or the screening of sequence libraries. Here, we present a simple generative approach for creating synthetic 5'-UTR libraries based solely on the genomic sequence statistics of any desired organism. The method uses a sliding-window n-gram language model applied to native 5'-UTR sequences to produce novel sequences that preserve organism-specific base distributions and motifs without hard-coding specific motifs or mechanistic rules into inflexible statistical templates. We have applied this approach to the model bacterium Escherichia coli and the non-model probiotic Limosilactobacillus reuteri. Libraries of approximately 1,000 sequences were generated for each organism, from which about 100 unique sequences were experimentally tested for translation of a fluorescent reporter protein. In both organisms, the synthetic libraries yielded a broad range of translation levels from this relatively small number of tested variants. Sequences derived from an organism's own genomic statistics provided a more uniformly distributed range of translation rates in that organism than sequences derived from the other species. Correlations of individual sequence performance across the two species were weak, and thermodynamic predictions of ribosome binding strength showed very little predictive power, especially in the non-model L. reuteri. The results demonstrate that simple statistical language model approaches applied to genomic data can generate functional translational regulatory sequence libraries without detailed mechanistic knowledge or explicit reference to consensus motifs. The approach requires minimal computational resources, avoids reproducing native sequences, and can be readily applied to any organism with a sequenced genome. This strategy may lower technical barriers to expression tuning in non-model organisms.

5' Untranslated Regions

Enzymatic depletion of transposable elements in sequencing libraries and its application for genotyping multiplexed CRISPR-edited plants.

Whole-genome sequencing has become a common strategy to genotype individual plants of interest. Although a limited number of genomic regions usually need to be surveyed with this strategy, excess sequencing information is almost always generated at an appreciable financial cost. Repetitive sequences (e.g., transposons), which can account for more than 80% of the genome of some plants, are often not required in these genotyping projects. Therefore, strategies that enrich DNA coding for the protein-coding genes prior to sequencing can lower the cost to obtain sufficient sequence information. Here, we present the development and application of methylation-sensitive reduced representation sequencing (MsRR-Seq), which relies on the cytosine methylation-sensitive restriction enzyme MspJI to deplete constitutive heterochromatic DNA before library construction. By applying MsRR-Seq to citrus and maize, we show that protein-coding genes can be enriched in sequencing datasets. We then describe the application of MsRR-Seq to facilitate the identification of complex mutants from populations of citrus plants resulting from multiplex CRISPR/Cas9 editing of four genes. Overall, this work demonstrates an easy and low-cost method to enrich non-repetitive DNA in high-throughput sequencing libraries, an approach that is especially useful for large plant genomes with an excessively high proportion of methylated repetitive sequences.

DNA Transposable Elements

IMPACT OF FLUORESCENT DYES ON MUTATIONS IN NEXT GENERATION SEQUENCING LIBRARY GENERATION.

DNA labelling fluorescent dyes such as ethidium bromide have long been considered to be highly mutagenic during DNA replication. While recent studies have pushed back on this narrative, the intercalative nature of these dyes continues to raise the possibility that these dyes can induce mutations. The iconPCR instrument by n6tec uses fluorescent dyes to measure amplification in real time and to adjust cycling conditions. However, since this use of qPCR is preparative and not analytical, mutations introduced by fluorescent dyes would be propagated into the sequencing reaction. To address the impact of these dyes on downstream analyses, we have performed routine mutation calling as well as mutational signature analysis on samples amplified using the iconPCR in the presence of either SYBR or EvaGreen. Sequence analysis revealed very minimal impacts of dyes on the reactions, largely within the noise regimen with only subtle changes in mutation rates seen. Mutational signature analysis was unable to identify any key signatures assignable to the dyes in either substitutions or indel domains. The mutational impact of intercalating dyes during fluorescence-guided amplification is therefore minimal and can be disregarded in all but the most sensitive NGS applications.

Fluorescent Dyes

A scalable, low-cost, sample hashing workflow for multiomic single-cell analysis using the Seq-Well S3 platform.

In-depth analyses of clinical samples have the potential to provide unparalleled insights into the cellular mechanisms that underlie both health and disease, as well as therapeutic and prophylactic responses. However, these specimens are often paucicellular, necessitating the use of workflows that maximize the amount of information that can be learned. Here we provide a detailed protocol for generating and analyzing single-cell multiomic data from low-input samples with the Seq-Well S3 platform. We further describe a matched pipeline for sample hashing that reduces costs and sources of technical variation in the resulting data while also enhancing throughput. In brief, our streamlined and efficient methodology involves: (1) optionally staining single-cell suspensions with antibody-oligonucleotide conjugates for cell surface protein quantification and/or sample multiplexing; (2) generating Seq-Well S3 sequencing libraries; (3) optionally producing bulk-RNA sequencing libraries via SMART-seq2 to support genetic demultiplexing; and (4) computationally analyzing the resulting data. Each step herein has been designed to leverage readily available reagents and standard laboratory equipment, substantially lowering barriers to entry for researchers. The overall Protocol can yield high-quality multiomic insights from samples in under a week.

Single-Cell Analysis

TGIRT-seq to profile tRNA-derived RNAs and associated RNA modifications.

RNA modifications are key regulators for RNA processes. tRNA-derived RNAs are small RNAs with size between 15 and 50 bases long that are processed from mature or precursor tRNAs. Despite their more recent discovery, tRNA-derived RNAs have been found to play regulatory roles in many cellular processes including gene silencing, protein synthesis, stress response, and transgenerational inheritance. Furthermore, tRNA-derived RNAs are highly abundant in bodily fluids, posing as potential biomarkers. A unique feature of tRNA-derived RNAs is that they are rich in RNA modifications. Many of the RNA modifications on tRNA-derived RNAs disrupt Watson-Crick base pairing and will thus stall reverse transcriptase, such as N1-methyladenosine (m1A), N1-methylguanosine (m1G) and N2, N2-dimethylguanosine (m22G). These RNA modifications add another layer of regulation onto tRNA-derived RNAs' functions and are of interests for future research. However, these RNA modifications could also lead to lower detection of modification-containing RNAs in genome-wide small RNA sequencing analysis due to reverse transcriptase stall. To circumvent this bias, TGIRT (Thermostable Group II Intron Reverse Transcriptase) has been used to readthrough RNA modifications inserting mismatches. These mismatch signatures can then be used to precisely map the modification sites at base resolution. Here we describe the step-by-step experimental protocol to start with purified RNAs from cells or tissues and use TGIRT to make small RNA sequencing library for Illumina sequencing to profile the abundance of tRNA-derived RNAs and the associated RNA modifications.

RNA, Transfer

Detection of mitochondrial tDRs in killifish embryos and other non-model organisms.

In recent years a diversity of small noncoding RNAs have been identified that originate from the mitochondrial genome. These mitosRNAs are often dominated by tRNA-derived small RNAs (mito-tDRs). Differential expression of mito-tDRs is associated with responses to stress. They also appear to be expressed differentially during development and their expression may be enriched in stress-tolerant animals. Very little is currently known about roles or modes of action of these sequences, although they are implicated in a diversity of processes such as cell cycle regulation, mRNA stability, regulation of ROS production, and import of proteins into the mitochondrion. To better understand the various roles these sequences may play, it is critical that we understand their diversity, cellular location, and the context for their expression. This protocol outlines the methodologies used to detect mitosRNAs, including mito-tDRs, in embryos and cells of the annual killifish Austrofundulus limnaeus. We highlight critical steps in the isolation of RNA, creation of sequencing libraries, bioinformatics processing of sequence data, and methods for validation of expression that support a robust discovery pipeline for mitosRNAs even from species with incomplete reference genome sequences.

Animals

QPL-enabled HTSlib library: accelerating sequence file compression using Intel IAA.

SUMMARY: Sequence analysis workflows require the accessibility of large datasets, which require state-of-the-art compression tools. These compression tools, such as Samtools, often rely on HTSlib as a GZip implementation, but are still limited by throughput on time-intensive compression. QPL-HTSLib offers order of magnitude speedups for the compression and decompression of the SAM and BAM file formats commonly used in genomics workflows at the cost of a slightly larger compressed file, and is a drop-in replacement for HTSlib on Intel systems. AVAILABILITY AND IMPLEMENTATION: QPL-HTSLib is freely available on Github as an open-source software project.

Journal Article

An enhanced multisegment RT-PCR method for influenza A virus sequencing: Improved performance and reduced preparation time over traditional methods.

Influenza A viruses (IAVs) remain a major global health threat, affecting both human and animal populations. Whole-genome sequencing is essential for monitoring viral evolution, zoonotic transmission, and emerging variants. However, conventional RT-PCR methods often result in incomplete gene coverage, amplification biases, and reduced sequencing accuracy, particularly in clinical samples. We developed a robust In-house method for IAV full-genome sequencing using the Oxford Nanopore Technologies (ONT) long-read sequencing platform. This method integrates an in-house multisegment Reverse Transcription PCR (RT-PCR) method with a streamlined 2-pool primer design targeting all eight IAV gene segments. RNA extracted from clinical and stock virus samples was reverse-transcribed and amplified using Superscript IV-based chemistry, followed by magnetic bead purification to ensure high-quality amplicons. Sequencing libraries were prepared with the Native Barcoding Kit 24 (SQK-NBD114.24) and sequenced on R10.4.1 flow cells on the MinION MK1C device. Data analysis using the Iterative Refinement Meta-Assembler (IRMA) confirmed improved read depth, uniform coverage, and complete genome recovery. Compared to conventional methods, our In-House Multisegment 2-Pool (IH-MS2P) RT-PCR method generated higher numbers of matched read counts, minimized chimeric artifacts, and delivered superior genome coverage across human, swine, and avian isolates. This optimized RT-PCR method provides a high-performance, time-efficient, and portable solution for influenza genomics, demonstrating robust applicability even with clinical samples of low RNA yield.

Influenza A virus

Classification and sequencing of hepatitis D virus from a large cohort of chronically infected individuals paired with co-infecting hepatitis B virus sequencing: a genomic characterisation study.

BACKGROUND: The most severe form of viral hepatitis is caused by co-infection of hepatitis D virus (HDV) and hepatitis B virus (HBV). Phylogenetic analyses classify HBV and HDV into eight major genotypes: HBV GTA to GTH and HDV GT1 to GT8. Paired HBV and HDV sequencing data from participants with chronic hepatitis delta are scarce. We aimed to sequence and genotype HDV and HBV from a large cohort of participants from clinical studies and diverse countries of origin. METHODS: 407 participants with chronic hepatitis D from 24 countries were characterised (124 participants from MYR301 clinical trial, 93 from MYR204, 114 from MYR202, and an additional 76 participants from diverse geographical locations). HBV and HDV from participants were analysed using sequencing, enzyme immunoassay, or both to determine HBV and HDV genotypes. BLAST analysis and phylogenetics were used to determine HBV and HDV genotypes with reference sequence libraries. Bulevirtide treatment response (measured by HDV RNA decline and normalisation of alanine aminotransferase) was compared by genotype for MYR trial participants. FINDINGS: HDV sequencing assays were successful for 386 (95%) of 407 participants and HBV sequencing or serology-based HBV genotyping assays were successful for genotyping 395 (97%) participants. For individual genotypes, HBV GTD (336 [83%] participants) and HDV GT1 (364 [89%]) were the most prevalent. For paired HBV-HDV genotypes, HBV-HDV D/1 was most common (320 [79%] of 407) followed by A/1 (30 [7%]). Phylogenetic analyses of HDV full-genome sequences showed distinct clusters of sequences within HDV GT1, and four novel provisional HDV GT1 subgenotypes, HDV GT1fp to HDVGT1ip, were identified. For 218 MYR clinical trial participants, bulevirtide treatment response was similar across HDV GT1 subgenotypes (both established and newly identified). INTERPRETATION: Novel HDV subgenotypes identified in this study indicate a greater genetic diversity of HDV GT1 than previously recognised. This knowledge will be important for developing better diagnostics, and in understanding HDV genotype-specific biology and response to treatment. More extensive HDV sequencing from under-sampled regions, such as Africa, is needed to determine the true breadth of HDV sequence and genotype diversity. FUNDING: Gilead Sciences.

Hepatitis Delta Virus

Volumetric DNA microscopy for mapping spatial transcriptomes in three dimensions.

The architecture and function of biological systems are inherently three-dimensional, yet most existing spatial transcriptomic technologies remain restricted to thin tissue sections, limiting their capacity to resolve cellular organization and microenvironments within intact tissue volumes. To address this limitation, we developed volumetric DNA microscopy, a scalable, optics-free approach for spatial transcriptome profiling directly within intact biological specimens. The method encodes spatial information into DNA molecules that form a dense intermolecular network in situ, enabling the reconstruction of three-dimensional spatial relationships through short-read sequencing and computational analysis. Here we detail the complete workflow including in situ cDNA synthesis, spatial encoding through DNA nanoball formation, dual-scale proximity bridging between neighboring nanoballs and spatial reconstruction via geodesic spectral embedding. Sequencing libraries can be generated within 7-8 d by a competent graduate-level molecular biologist, followed by standardized downstream computational analysis. Because the workflow requires only routine molecular biology reagents and a benchtop sequencer, volumetric DNA microscopy provides a versatile platform for exploring genetic and morphological features in intact tissues.

Spatial Transcriptomics

Comprehensive evaluation of new sequencer T20 and well-established T7 with 507 human samples.

The DNBSEQ-T20×2 (T20) sequencer, developed by MGI Tech, enables cost-effective human whole-genome sequencing (WGS) at 30× coverage for less than $100 per genome. Here, we evaluate the sequencing performance and data quality of the T20 platform by benchmarking it against the established DNBSEQ-T7 (T7) sequencer using 507 samples derived from blood (N = 75), stool (N = 242), and saliva (N = 190). The T20 exhibited lower sequencing quality metrics compared with the T7, with Q20 scores of 95.76%-95.83% and Q30 scores of 87.25%-87.40%, compared with 97.81%-97.93% and 93.26%-93.60%, respectively, for T7 data. Quality differences were more evident toward the end of reads, and PCR-free libraries sequenced on the T20 showed similar reductions in quality scores. The median empirical base error rate estimated from 102 ZymoBIOMICS samples was 0.33%. The T20 demonstrated comparable coverage uniformity to the T7 and showed high concordance in microbiome composition analysis, with a median Bray-Curtis dissimilarity of 0.02. Variant calling performance was highly consistent between the two platforms. Among variants with non-missing genotype calls on both platforms, 94.92% of SNPs and 87.20% of InDels showed concordant genotypes between T20 and T7. Overall, the T20 delivers reliable sequencing accuracy and reproducibility for large-scale genomic and microbiome studies, providing a cost-effective alternative for high-throughput sequencing applications.

Metagenomics

Tracking-seq: a universal off-target detection approach for CRISPR-Cas genome editing.

Tracking-seq is a highly sensitive method for genome-wide detection of off-target effects in cells edited with diverse genome editing modalities, including Cas9, cytosine base editors, adenine base editors and prime editors. Since most genome editors induce DNA repair pathways and generate single-stranded DNA (ssDNA) intermediates, Tracking-seq leverages this process by tracking replication protein A-a key protein that binds and protects ssDNA-to identify on-target and off-target events. Here we provide a detailed protocol for Tracking-seq, covering genome editing of cells, extraction of replication protein A-bound ssDNA, sequencing library construction and data analysis using our custom computational tool Offtracker. Tracking-seq is applicable to various genome editing scenarios with low cell input, delivering high-performance results. The entire workflow, from genome editing to data analysis, can be completed within 1-2 weeks, making it a rapid solution for assessing genome-wide off-target activity.

CRISPR-Cas Systems

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

Construction of a prognostic model for gastric cancer based on immune infiltration and microenvironment, and exploration of MEF2C gene function.

BACKGROUND: Advanced gastric cancer (GC) exhibits a high recurrence rate and a dismal prognosis. Myocyte enhancer factor 2c (MEF2C) was found to contribute to the development of various types of cancer. Therefore, our aim is to develop a prognostic model that predicts the prognosis of GC patients and initially explore the role of MEF2C in immunotherapy for GC. METHODS: Transcriptome sequence data of GC was obtained from The Cancer Genome Atlas (TCGA), the Gene Expression Omnibus (GEO) and PRJEB25780 cohort for subsequent immune infiltration analysis, immune microenvironment analysis, consensus clustering analysis and feature selection for definition and classification of gene M and N. Principal component analysis (PCA) modeling was performed based on gene M and N for the calculation of immune checkpoint inhibitor (ICI) Score. Then, a Nomogram was constructed and evaluated for predicting the prognosis of GC patients, based on univariate and multivariate Cox regression. Functional enrichment analysis was performed to initially investigate the potential biological mechanisms. Through Genomics of Drug Sensitivity in Cancer (GDSC) dataset, the estimated IC50 values of several chemotherapeutic drugs were calculated. Tumor-related transcription factors (TFs) were retrieved from the Cistrome Cancer database and utilized our model to screen these TFs, and weighted correlation network analysis (WGCNA) was performed to identify transcription factors strongly associated with immunotherapy in GC. Finally, 10 patients with advanced GC were enrolled from Sun Yat-sen University Cancer Center, including paired tumor tissues, paracancerous tissues and peritoneal metastases, for preparing sequencing library, in order to perform external validation. RESULTS: Lower ICI Score was correlated with improved prognosis in both the training and validation cohorts. First, lower mutant-allele tumor heterogeneity (MATH) was associated with lower ICI Score, and those GC patients with lower MATH and lower ICI Score had the best prognosis. Second, regardless of the T or N staging, the low ICI Score group had significantly higher overall survival (OS) compared to the high ICI Score group. For its mechanisms, consistently, for Camptothecin, Doxorubicin, Mitomycin, Docetaxel, Cisplatin, Vinblastine, Sorafenib and Paclitaxel, all of the IC50 values were significantly lower in the low ICI Score group compared to the high ICI Score group. As a result, based on univariate and multivariate Cox regression, ICI Score was considered to be an independent prognostic factor for GC. And our Nomogram showed good agreement between predicted and actual probabilities. Based on CIBERSORT deconvolution analysis, there was difference of immune cell composition found between high and low ICI Score groups, probably affecting the efficacy of immunotherapy. Then, MEF2C, a tumor-related transcription factor, was screened out by WGCNA analysis. Higher MEF2C expression is significantly correlated with a worse OS. Moreover, its higher expression is also negatively correlated with tumor mutation burden (TMB) and microsatellite instability (MSI), but positively correlated with several immunosuppressive molecules, indicating MEF2C may exert its influence on tumor development by upregulating immunosuppressive molecules. Finally, based on transcriptome sequencing data on 10 paired tumor tissues from Sun Yat-sen University Cancer Center, MEF2C expression was significantly lower in paracancerous tissues compared to tumor tissues and peritoneal metastases, and it was also lower in tumor tissues compared to peritoneal metastases, indicating a potential positive association between MEF2C expression and tumor invasiveness. CONCLUSIONS: Our prognostic model can effectively predict outcomes and facilitate stratification GC patients, offering valuable insights for clinical decision-making. The identified transcription factor MEF2C can serve as a biomarker for assessing the efficacy of immunotherapy for GC.

Humans

Atherosclerotic plaque fibroblasts derive from adventitial and medial Pdgfra-lineage-positive cells and predominantly maintain fibroblast identity.

AIMS: Fibroblasts are mesenchymal cells in the healthy vascular adventitia. In atherosclerosis, single-cell sequencing datasets suggest fibroblasts are abundant in plaques. However, their identity, origin, and fate during plaque progression remain unclear, which we aim to unravel here. APPROACH AND RESULTS: To robustly define fibroblast identity, origin, and fate, we employed meta-analyses of 54 single-cell RNA sequencing libraries, including murine smooth muscle cell (Myh11) and endothelial cell (EC) (Cdh5) lineage reporter mice with and without atherosclerosis; human control and atherosclerotic arteries; and murine adventitia and atherosclerotic plaques processed separately from low-density lipoprotein (LDL) receptor knockout (Ldlr-/-) mice. These meta-analyses showed that murine and human plaque fibroblast identity was robustly defined by Pdgfra, Pi16, Cygb, and Serpinf1 mRNA. Ninety-five percent of plaque fibroblasts do not derive from the Myh11 lineage, while no Cdh5-lineage-positive cells were present in the fibroblast cluster. We identified five murine arterial fibroblast subsets in atherosclerotic murine aorta: progenitor fibroblasts, matrix fibroblasts, inflammatory fibroblasts, an EC-like fibroblast subset, detected in both adventitia and plaques, and Col5a3+ fibroblasts, unique to the adventitia. We next studied fibroblast identity, origin, and fate using pseudotime analysis and Pdgfra-CreERT2/tdTomato lineage reporter mice (Pdgfra Lin+). Healthy Pdgfra Lin+ reporter mice showed predominant adventitial tdTomato expression, and infrequent medial and intimal Pdgfra Lin+ cells co-expressing MYH11 and PECAM1, respectively. The Pdgfra Lin+ plaque area increased with diet duration. Pdgfra Lin+ cells largely maintain fibroblast identity in the plaque, while <10% co-express SMC markers (MYH11, SM22&#x3b1;), or contribute to ACTA2+ cap cells. ECs gaining mesenchymal markers are transcriptionally distinct from Cdh5-lineage-negative fibroblasts gaining EC markers. Plaque-resident EC-like fibroblasts displayed a mesenchymal-to-endothelial transition transcriptome, which was induced in human primary fibroblasts in vitro by starvation, and dampened or reversed by IL1B, TGFB1, TGFB3, and oxidized LDL. Cross-species integration showed that all murine plaque fibroblasts were conserved in human atherosclerosis, with one additional subset partially resembling murine subsets, and three human-specific subsets. Importantly, human fibroblast subsets differentially correlated to human plaque traits, with EC-like fibroblasts correlating to plaque instability. CONCLUSION: Our results indicate that 95% of plaque-residing fibroblasts are Myh11 Lin- Plaque fibroblasts have a dual origin, predominantly adventitial Pdgfra Lin+ progenitor fibroblasts, with a minor contribution from medial Pdgfra Lin+ &#xa0;Myh11+ SMCs. Most plaque fibroblasts maintain fibroblast identity. Murine plaque fibroblast subsets were conserved in human atherosclerosis. EC-like fibroblasts are linked to human plaque instability. Intervening in progenitor-to-specific fibroblast transitions could present a new avenue to promote plaque stability in atherosclerosis.

Atherosclerosis

Uncovering host transcriptional responses to tilapia lake virus (TiLV) through De novo RNA-seq assembly in Nile tilapia, Oreochromis niloticus.

Tilapia lake virus (TiLV) has emerged as an important pathogen that negatively impacts tilapia farming globally. Using RNA sequencing technology, this study investigated the liver transcriptomic profile of apparently healthy and TiLV-infected Oreochromis niloticus from wild. RNA sequence libraries generated 3,356 differentially expressed genes (DEGs), with 1,726 genes that were upregulated. Kyoto Encyclopedia of Genes and Genomes (KEGG) analysis identified 680 different pathways with differential regulation of the metabolic and immune-related pathways indicating that TiLV may interfere in host metabolism and replicate to establish the infection. This study provides transcriptomic insights into the liver responses of naturally TiLV-infected wild O. niloticus and highlights key immune and metabolic pathways associated with viral infection.

Animals

CRISPGen: A deep generative framework for multi-objective CRISPR/Cas9 guide RNA design via Conditional Latent Diffusion and Dual-Critic Reinforcement Learning.

MOTIVATION: The CRISPR-Cas9 system offers transformative potential for precision genome editing, yet its clinical translation remains constrained by the risk of unintended off-target double-strand breaks. While current discriminative models excel at evaluating pre-specified candidate guides, resolving the fundamental antagonism between on-target cleavage efficiency and off-target specificity within a fixed sequence search space remains a major challenge. RESULTS: We present CRISPGen, a unified deep generative framework that reframes sgRNA design as a multi-objective constrained sequence synthesis problem. It integrates (i) DNABERT-2 genomic-language embeddings, (ii) a conditional latent diffusion generator conditioned on a user-specified on-target efficiency target, and (iii) a dual-critic reinforcement-learning (RL) stage that couples a frozen on-target efficiency critic with a cross-attention off-target discriminator (validation Pearson R=0.8157) trained on a unified corpus of experimental off-target events from six detection platforms. Across 1000 generated sgRNAs, CRISPGen reduces the mean off-target discriminator score by 99.7% relative to the pre-RL baseline and, under an exhaustive whole-genome screen of all 302,631,056 NGG PAM sites in GRCh38, yields zero perfect-match and only 55 one-mismatch genomic hits. We further show, transparently, that the internal on-target critic saturates under RL optimization - an instance of Goodhart's Law - and therefore assess on-target viability using an independent external CRISPRon screen (mean 47.10/100). Repeating the RL fine-tuning stage under three random seeds (with the diffusion generator, DNABERT-2 embeddings, and off-target discriminator held fixed) yields a stable operating point across seeds. Full diversity, per-mismatch, and reproducibility statistics are reported in the Results. AVAILABILITY: Source code is available at https://github.com/malekpouri/CRISPGen; the pre-trained checkpoints and the 3,000,000-sequence library are hosted on Hugging Face (https://huggingface.co/malekpouri/CRISPGen-Checkpoints) and archived on Zenodo under DOI 10.5281/zenodo.21428641.

CRISPR-Cas9

Integrative dual-track transcriptomics reveals stage-specific coordination, regulatory divergence, and HSP90AA1-associated remodeling in human folliculogenesis.

Human folliculogenesis depends on coordinated yet non-identical developmental remodeling in the oocyte and its surrounding granulosa cells. When these two compartments remain synchronized and when they diverge into lineage-specific regulatory states, however, remains incompletely resolved. Here we performed an integrative dual-track re-analysis of the human RNA-seq dataset GSE107746, modeling oocytes and granulosa cells as distinct but developmentally linked compartments across follicular progression. Analysis of 148 sequencing libraries showed that compartment identity was the dominant source of transcriptomic variation, supporting compartment-aware downstream interpretation. Within this framework, oocytes followed a relatively continuous developmental trajectory, with substantial transcriptional remodeling already evident across adjacent stages, whereas granulosa cells showed weaker early-stage contrasts but markedly stronger late-stage reorganization, particularly around the antral and preovulatory transitions. Functional enrichment indicated that oocyte maturation was associated with RNA-processing and broader genome-regulatory remodeling, whereas granulosa maturation was dominated by progressive mitochondrial and bioenergetic activation. Co-expression analysis showed that both compartments contained strong late-stage programmes together with inverse early-state modules, indicating a shared systems-level architecture of maturation, although the hub-gene composition and biological content of these programmes were largely compartment-specific. Machine-learning validation reinforced this asymmetry: oocyte stage classification was best recovered from a compact eigengene-based representation, whereas granulosa stage discrimination was better resolved by a broader differential-expression-derived feature set. At the gene level, HSP90AA1 emerged as a stage-associated marker with compartment-specific behavior, showing progressive attenuation across oocyte development, assignment to the selected oocyte blue module, and sharper transitional dynamics in granulosa cells. Together, these findings support a model in which human folliculogenesis proceeds through coordinated but non-equivalent transcriptomic remodeling, with shared developmental logic at the systems level but distinct molecular execution in germline and somatic compartments.

Co-expression networks