PubMed HealthSearch

SEARCH · PubMed Health

Results for “Nucleotide signature”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Predicting host tropism in influenza a viruses: insights from multi-segment nucleotide signatures.

BACKGROUND: Influenza A virus (IAV) poses a significant public health threat due to its cross-species transmission and complex host adaptation mechanisms. This study integrated whole-genome data from avian, human, swine, and bovine IAV strains, using machine learning to predict viral host tropism based on nucleotide site features and to identify key sites driving host adaptation along with their synergistic effects. METHODS: A total of 64,000 IAV sequences from avian, human, swine, and bovine hosts were analyzed to build host-prediction models. A four-class classification framework (avian, human, swine, bovine) was constructed using nucleotide site features from all eight genomic segments (PB2, PB1, PA, HA, NP, NA, MP, NS). Eight machine learning algorithms (logistic regression, decision tree, random forest, SVM, KNN, gradient boosting, XGBoost, LightGBM) were benchmarked via 10-fold stratified cross-validation. Model performance was evaluated using accuracy, precision, recall, F1-score, AUPRC, and AUC. SHAP (SHapley Additive exPlanations) analysis prioritized critical nucleotide sites, while bivariate association tests identified synergistic/antagonistic interactions between sites. Nucleotide composition profiles were compared across host groups using hierarchical clustering and heatmap visualization. RESULTS: The XGBoost algorithm demonstrated the best and most stable performance, achieving an AUC value of over 0.95 in distinguishing human-derived sequences from non-human ones. SHAP analysis identified the top 20 critical nucleotide sites for each gene segment, such as sites 46 and 698 in the NS segment. Nucleotide composition analysis revealed high similarity between human and swine sequences in the HA and PB2 segments, and between avian and bovine sequences. The HA segment was particularly challenging in differentiating human from swine strains. Bivariate site association analysis uncovered significant synergistic or antagonistic effects between key sites within gene segments, forming complex networks. For instance, in the NS segment, a positive prediction contribution was observed when sites 371, 698, and 419 were all G. CONCLUSIONS: This study advances our mechanistic understanding of IAV host adaptation, identifies molecular determinants for zoonotic risk stratification, and establishes a scalable machine learning framework for predicting viral host tropism through nucleotide signature analysis, thereby enhancing surveillance strategies and informing preventive measures against emerging viral threats.

Influenza A virus

When light colour matters: Spectral quality is associated with distinct small RNA candidates in Arabidopsis thaliana.

The spectral composition of incident light is perceived by plant photoreceptors and can rapidly reshape downstream gene expression programs; however, its impact on the small-RNA layer beyond annotated miRNAs remains incompletely characterized. The objective of this study was to determine whether 3 h exposure of Arabidopsis thaliana rosettes to blue, green, red, or white light at an equal photosynthetic photon flux density (PPFD) of 400 μmol m-2 s-1 is associated with distinct profiles of candidate small RNAs. A stringent discovery and post-processing workflow was applied to identify high-confidence, treatment-associated small RNA candidates beyond annotated miRNA. The miRkwood-based pipeline, combined with additional filtering and contextual annotation, recovered a compact set of candidates dominated by the blue-light treatment (41 candidates), compared with fewer candidates associated with red (8), green (6), and white light (3). Genomic-context analysis indicated that most candidate sites were intergenic, with candidates detected under green light being entirely intergenic, and overlap with transposon annotations was used to distinguish candidates arising from transposon-rich genomic space. Sequence-feature profiling revealed pronounced treatment-dependent terminal nucleotide biases at both 5' and 3' ends, consistent with spectrum-associated shifts in length-class and terminal-nucleotide signatures that are informative for AGO-loading hypotheses. Target prediction highlighted a subset of genes showing convergent targeting by multiple independent blue-associated candidates, and duplex summaries showed structured, plant-like complementarity patterns (including frequent 10-11 pairing). Together, these results indicate that 3 h exposure to wavelength-defined light at 400 μmol m-2 s-1, particularly blue light, is associated with a distinct profile of detectable candidate small RNAs in Arabidopsis leaves and identify candidate interactions for follow-up validation.

5′ nucleotide bias

Write and Read: Harnessing Synthetic DNA Modifications for Nanopore Sequencing.

An exciting feature of nanopore sequencing is its ability to record multi-omic information on the same sequenced DNA molecule. Well-trained models allow the detection of nucleotide-specific molecular signatures through changes in ionic current as DNA molecules translocate through the nanopore. Thus, naturally occurring DNA modifications, such as DNA methylation and hydroxymethylation, may be recorded simultaneously with the genetic sequence. Additional genomic information, such as chromatin state or the locations of bound transcription factors, may also be recorded if their locations are chemically encoded into the DNA. Here, we present a versatile "write-and-read" framework, where chemo-enzymatic DNA labeling with unnatural synthetic tags results in predictable electrical fingerprints in nanopore sequencing. As a proof-of-concept, we explore a DNA glucosylation approach that selectively modifies 5-hydroxymethylcytosine (5hmC) with glucose or glucose-azide adducts. We demonstrate that these modifications generate distinct and reproducible electrical shifts, enabling the direct detection of chemically altered nucleotides. We further demonstrate that enzymatic alkylation, such as the enzymatic transfer of azide residues to the N6 position of adenines, also produces characteristic nanopore signal shifts relative to the native adenine and 6-methyladenine. Beyond direct nucleotide detection, this approach introduces new possibilities for bio-orthogonal DNA labeling, enabling an extended alphabet of sequence-specific detectable moieties. The future use of programmable chemical modifications for simultaneous analysis of multiple omics features on individual molecules opens new avenues for genetic research and discovery.

5-hydroxymethylcytosine (5hmC)

Whole-genome sequencing identifies genetic diversity and adaptive signatures of hypoxia and ultraviolet radiation in Chinese chickens.

INTRODUCTION: Domestic chickens primarily descended from the wild red junglefowl, play a crucial role in global egg and meat production. China hosts diverse indigenous chicken populations that have adapted to various environmental conditions, including high-altitude with hypoxic and ultraviolet radiation stress. METHOD: We analyzed whole-genome sequences of 118 birds from five Indigenous Chinese chicken populations and 295 chicken genomes from publicly available databases to identify genomic diversity, admixture, and selection signatures of chickens adapted to high-altitude environments. Selection signatures were identified using nucleotide diversity (π), Tajima's D, XPEHH, and XP-CLR, selection scan methods. RESULTS: We observed a reduction in genetic diversity and historical declines in effective population size in high-altitude chicken, suggesting ongoing selection pressures shaping these populations. Selection scans identified nine genomic regions under strong positive selection, enriched for genes associated with hypoxia and ultraviolet radiation. Notably, five genes (TPK1, BAZ2B, MARCHF7, LLGL2, and RCAN3) were repeatedly detected across multiple selection signature analyses. RNA-seq analysis further confirmed the differential expression of these genes in the lung and heart tissues of chickens adapted to high and low altitudes, reinforcing their role in physiological adaptation to hypoxic environments. Altitude adaptation is driven by the selection of genes involved in oxygen metabolism, cellular stress response, and energy regulation. CONCLUSION: Our study provides compelling genetic evidence for differentiation between high and low and high-altitude Chinese chicken populations. These findings also ensure our understanding of local adaptation in poultry and establish a genomic framework for breeding strategies to improve environmental resilience to altitude-related stressors.

Animals

Screening of the key single nucleotide polymorphisms in type 2 diabetes mellitus complicated with lower extremity arterial disease by machine learning.

OBJECTIVES: Diabetic lower extremity arterial disease (LEAD) is a manifestation of diabetic lower extremity vascular complications. This study aimed to screen the key single nucleotide polymorphism (SNP) gene signature in patients with type 2 diabetes mellitus (T2DM) and LEAD. METHODS: A total of 147 patients with T2DM complicated by LEAD and 144 patients with T2DM without LEAD were enrolled for transcriptome sequencing. The Plink software was used to preprocess the data. Five machine learning methods were adopted to build the SNP diagnosis models. The receiver operating characteristic (ROC) curve was used to quantify the predicted probabilities of the model. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed using the cluster Profiler package. Finally, regression statistical analysis was used to correlate the key SNPs with clinical information and biochemical indicators. RESULTS: A total of 24 SNPs were retained and 10 SNPs were risk allele genes. Nine SNPs (rs7412, rs1800629, rs699947, rs3918242, rs668, rs1800470, rs1800449, rs1800469, and rs1024611) were identified as the key SNPs sites. GO and KEGG pathway analyses revealed that these genes are mainly enriched in fluid shear stress and atherosclerosis. Finally, rs1800449 was associated with low-density lipoprotein cholesterol (LDL-C). With high density lipoprotein cholesterol (HDL-C), related site was rs1024611. The sites associated with total cholesterol (CHOL) were rs1800449 and rs7412.The site associated with apolipoprotein B (APOB) and apolipoprotein A1 (APOA1) were rs1800470 and rs1800469. CONCLUSION: This study authenticated nine SNPs for the diagnosis of T2DM patients with LEAD, which will be of great significance in the development of diagnostic molecular biomarkers for T2DM patients.

Humans

TSC angiofibroma and ungual fibroma have different mutation signatures, with recurrent mutations in KMT2C.

PURPOSE: Tuberous sclerosis complex (TSC) is an autosomal dominant tumor suppressor syndrome characterized by tumors affecting multiple tissues, including skin, due to inactivating TSC1/TSC2 variants. Genome-wide profiling of somatic mutations in a unique collection of angiofibroma (FAF) and ungual fibroma (UF) TSC skin tumors was performed. METHODS: Genome sequencing was performed on 9 samples, comprising 4 FAF and 5 UF, along with 6 matched normal samples from 6 individuals with TSC. RESULTS: TSC-FAF and TSC-UF skin tumors have different mutation signatures, with a predominance of UV-related single-nucleotide variant (SNV; SBS7a and SBS7b) and dinucleotide variant (DNV; DBS1) signatures in FAF, and aging-related SNV (SBS1 and SBS5) signatures in UF. We also identified a novel DNV signature for TSC-UF, with frequent TG>CA and TT>GG substitutions. Furthermore, 3 inactivating somatic mutations in KMT2C were observed in 2 of 4 TSC-FAF and 5 mutations in other cancer genes. CONCLUSION: The distinct SNV mutation signatures seen in TSC-FAF and UF indicate that they develop through distinct pathogenic mechanisms, UV-induced mutagenesis in FAF, and aging-related mutagenesis in UF. The mechanism of the novel DNV signature in UFs merits further investigation. Our observation on the occurrence of KMT2C mutations suggests that KMT2C inactivation contributes to the pathogenesis of TSC-FAF.

Humans

Intra-colony divergence and global allele sharing reflect purifying selection and recombination at the Botryllus histocompatibility factor locus.

Urochordates, the closest relatives of vertebrates, lack adaptive immunity. However, some taxa, such as the colonial species Botryllus schlosseri, provide a unique model for studying innate self/non-self recognition through natural allogeneic transplantation responses. In this species, interactions between colonies are controlled by a highly polymorphic locus, with the Botryllus histocompatibility factor (BHF) being the only gene known to predict tissue fusion or rejection outcomes with complete accuracy. Here, we analyzed full-length BHF alleles from 19 laboratory-born and wild colonies and found that highly divergent alleles tend to coexist within individuals, whereas identical alleles can be shared across continental-scale distances. Despite extensive length variation, evidence of intragenic recombination, and pronounced nucleotide diversity, BHF exhibits limited protein divergence, with 33 alleles encoding only 17 distinct protein variants. Contrary to expectations for polymorphic recognition genes, no evidence of balancing or directional selection was detected. Instead, signatures of purifying selection were observed. We propose that this contrast between nucleotide and protein diversity arises from the combined effects of recombination, human-mediated gene flow, and linkage to nearby loci under balancing selection, while functional constraints maintain protein stability. These findings suggest that extensive protein diversification may not be a universal driver of allorecognition gene evolution.

Animals

Whole genome-based reclassification of the genus Metabacillus: Proposal for five novel genera, Chryseobacillus gen. nov., Cohnibacillus gen. nov., Salimetabacillus gen. nov., Pantoeobacillus gen. nov., and Lutimetabacillus gen. nov. and the description of one novel bacterial species, Chryseobacillus diguaensis sp. nov. isolated from soil in the Digua reservoir.

Comprehensive phylogenomic and comparative genomic analyses were conducted to clarify the taxonomic boundaries of the genus Metabacillus. Phylogenetic trees reconstructed from a set of single-copy orthologous proteins (SCOPs) revealed that the genus, as currently defined, is polyphyletic. The type species of the genus Metabacillus and its closest relatives formed a consistent clade, herein designated as Metabacillus sensu stricto. The remaining species were grouped into three well-supported clades: Kandeliae, Indicus, and Mangrovi, and two single-taxon lineages: M. arenae and M. lacus. The phylogenomic delineation found in these divergent taxa was corroborated by either inconsistent distribution patterns or the absence of previously defined conserved signature indels (CSIs) specific to Metabacillus. Genomic metrics, including Average Nucleotide Identity (ANI), Average Amino acid Identity (AAI), and digital DNA-DNA hybridization (dDDH) further supported the taxonomic delineation proposed here. The observed genomic divergence was mirrored by phenotypic differences, including variations in GC content ranges. Based on this polyphasic evidence, we propose the reclassification of the genus Metabacillus taxa into five novel genera: Chryseobacillus gen. nov. (encompassing the Kandeliae clade), Cohnibacillus gen. nov. (M. lacus), Salimetabacillus gen. nov. (M. arenae), Pantoeobacillus gen. nov. (Indicus clade), and Lutimetabacillus gen. nov. (Mangrovi clade). The core lineage is retained as Metabacillus sensu stricto, for which an emended description of the genus Metabacillus is also provided. A novel bacterial strain, designated as MAU-250T, was isolated from a soil sample collected on the shore of an artificial reservoir in the Andean foothills of the Maule Region in central Chile. Public metagenome screening supported a low-abundance taxon with broad ecological adaptability, preferentially associated with soil habitats. A polyphasic analysis based on phenotypic traits and genomic distances (78.0% ANIb and 19.8% dDDH against its closest relative) also supported its designation as a novel species, for which the name Chryseobacillus diguaensis sp. nov. is proposed. The type strain is MAU-250T (=RGM 3146T = IMI 507634T).

Phylogeny

Sequencing the orthologs of human autosomal forensic short tandem repeats provides individual- and species-level identification in African great apes.

BACKGROUND: Great apes are a global conservation concern, with anthropogenic pressures threatening their survival. Genetic analysis can be used to assess the effects of reduced population sizes and the effectiveness of conservation measures. In humans, autosomal short tandem repeats (aSTRs) are widely used in population genetics and for forensic individual identification and kinship testing. Traditionally, genotyping is length-based via capillary electrophoresis (CE), but there is an increasing move to direct analysis by massively parallel sequencing (MPS). An example is the ForenSeq DNA Signature Prep Kit, which amplifies multiple loci including 27 aSTRs, prior to sequencing via Illumina technology. Here we assess the applicability of this human-based kit in African great apes. We ask whether cross-species genotyping of the orthologs of these loci can provide both individual and (sub)species identification. RESULTS: The ForenSeq kit was used to amplify and sequence aSTRs in 52 individuals (14 chimpanzees; 4 bonobos; 16 western lowland, 6 eastern lowland, and 12 mountain gorillas). The orthologs of 24/27 human aSTRs amplified across species, and a core set of thirteen loci could be genotyped in all individuals. Genotypes were individually and (sub)species identifying. Both allelic diversity and the power to discriminate (sub)species were greater when considering STR sequences rather than allele lengths. Comparing human and African great-ape STR sequences with an orangutan outgroup showed general conservation of repeat types and allele size ranges. Variation in repeat array structures and a weak relationship with the known phylogeny suggests stochastic origins of mutations giving rise to diverse imperfect repeat arrays. Interruptions within long repeat arrays in African great apes do not appear to reduce allelic diversity. CONCLUSIONS: Orthologs of most human aSTRs in the ForenSeq DNA Signature Prep Kit can be analysed in African great apes. Primer redesign would reduce observed variability in amplification across some loci. MPS of the orthologs of human loci provides better resolution for both individual and (sub)species identification in great apes than standard CE-based approaches, and has the further advantage that there is no need to limit the number and size ranges of analysed loci.

Animals

Recurrent patterns of TOP1-mediated neuronal genomic damage shared by major neurodegenerative disorders.

Amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), and Alzheimer's disease (AD) represent two major categories of neurodegenerative disorders-TAR DNA-binding protein 43 (TDP-43) and tau proteinopathies-for which the mechanisms driving neuronal death remain unclear. Single-cell whole-genome sequencing of 469 neurons from C9ORF72 ALS, C9ORF72 FTD, AD, and control brains revealed increased somatic single-nucleotide variants (sSNVs) and insertions/deletions (sIndels) in all three diseases. Mutational signature analysis identified a disease-associated sSNV signature consistent with oxidative damage and an sIndel process affecting 22% of ALS, 76% of FTD, and 61% of AD neurons-but only 2% of control neurons-resembling signature ID4, previously linked to topoisomerase 1 (TOP1)-mediated mutagenesis. Rapid approach to DNA adduct recovery (RADAR) assays confirmed increased TOP1-DNA covalent complexes, and duplex sequencing confirmed the increased sIndels and identified single-strand events as likely precursor lesions. TOP1-associated sIndel mutagenesis and genome instability thus represent a mechanism shared by both TDP-43 and tau neurodegeneration.

Humans

Recurrent patterns of widespread neuronal genomic damage shared by major neurodegenerative disorders.

Amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), and Alzheimer's disease (AD) are common neurodegenerative disorders for which the mechanisms driving neuronal death remain unclear. Single-cell whole-genome sequencing of 429 neurons from three C9ORF72 ALS, six C9ORF72 FTD, seven AD, and twenty-three neurotypical control brains revealed significantly increased burdens in somatic single nucleotide variant (sSNV) and insertion/deletion (sIndel) in all three disease conditions. Mutational signature analysis identified a disease-associated sSNV signature suggestive of oxidative damage and an sIndel process, affecting 28% of ALS, 79% of FTD, and 65% of AD neurons but only 5% of control neurons (diseased vs. control: OR=31.20, p = 2.35×10-10). Disease-associated sIndels were primarily two-basepair deletions resembling signature ID4, which was previously linked to topoisomerase 1 (TOP1)-mediated mutagenesis. Duplex sequencing confirmed the presence of sIndels and identified similar single-strand events as potential precursor lesions. TOP1-associated sIndel mutagenesis and resulting genome instability may thus represent a common mechanism of neurodegeneration.

Journal Article

Understanding disease-associated metabolic changes in human colonic epithelial cells using the iColonEpithelium metabolic reconstruction.

The colonic epithelium plays a key role in the host-microbiome interactions, allowing uptake of various nutrients and driving important metabolic processes. To unravel detailed metabolic activities in the human colonic epithelium, our present study focuses on the generation of the first cell-type-specific genome-scale metabolic model (GEM) of human colonic epithelial cells, named iColonEpithelium. GEMs are powerful tools for exploring reactions and metabolites at the systems level and predicting the flux distributions at steady state. Our cell-type-specific iColonEpithelium metabolic reconstruction captures genes specifically expressed in the human colonic epithelial cells. iColonEpithelium is also capable of performing metabolic tasks specific to the colonic epithelium. A unique transport reaction compartment has been included to allow for the simulation of metabolic interactions with the gut microbiome. We used iColonEpithelium to identify metabolic signatures associated with inflammatory bowel disease. We used single-cell RNA sequencing data from Crohn's Diseases (CD) and ulcerative colitis (UC) samples to build disease-specific iColonEpithelium metabolic networks in order to predict metabolic signatures of colonocytes in both healthy and disease states. We identified reactions in nucleotide interconversion, fatty acid synthesis and tryptophan metabolism were differentially regulated in CD and UC conditions, relative to healthy control, which were in accordance with experimental results. The iColonEpithelium metabolic network can be used to identify mechanisms at the cellular level, and we show an initial proof-of-concept for how our tool can be leveraged to explore the metabolic interactions between host and gut microbiota.

Humans

Genome-wide scan for selection signatures in Mexican Sardo Negro Zebu cattle.

The Sardo Negro cattle (SN) is the only zebu cattle breed developed in Mexico. Since its development, the selection could have led to an increase in the homozygosity level in some regions of the genome and made differentiation with other cattle populations. We aimed to identify and characterize selection signatures in SN using medium-density SNP data using four approaches: 1) Runs of homozygosity (ROH) 2) Nucleotide Diversity 3) Tajima's D and 4) the Wright's fixation index (FST). A sample of 555 SN animals genotyped for 65k SNPs was used to obtain ROH segments considered regions under selection. The FST values were estimated by comparing the sample of genotyped SN animals with samples of genotyped animals from the Gir, Brahman, and Ongole breeds. Only one region mapped to 35.78-42.51 Mb on BTA6 was considered a selection signature by the ROH method. This selection signature overlapped with the lowest diversity, negative values of Tajima's D and a diversification region between SN and the other Zebu breeds by FST. We found several candidate genes (LCORL, NCAPG, and SLIT2) related to growth and other economically important productive traits in this common region. Using the FST method, different regions, such as regions on BTA8 (8:93.4-93.9 Mb), BTA11 (11:99.2-99.7), and BTA14 (14: 26.1-26.8) related to growth and milk traits also were defined as candidate selection signatures. The selective signals identified in this study reflected the direction of the selection pressure that primarily involves the increase of live weight traits in the Sardo Negro cattle breeding program.

Animals

Genetic diversity of Plasmodium falciparum helical interspersed subtelomeric (phistb) gene in Tanzania and neighboring countries.

BACKGROUND: Lysine-rich membrane associated Plasmodium helical interspersed subtelomeric gene (phistb) is a member of the phist family of genes which encodes exported proteins essential for the parasite's survival within infected red blood cells. Recent studies suggest the phistb gene as a promising malaria vaccine candidate, however, its genetic diversity remains understudied. This study assessed the genetic diversity of the phistb gene in regions of varying malaria transmission aiming to generate data and improve our understanding of this promising malaria vaccine candidate gene. METHODS: Genomic data from 1472 Plasmodium falciparum samples from Tanzania, Kenya, Uganda, and Ethiopia were retrieved in variant Calling file format (VCF) format from the MalariaGEN Pf7 database. Variants were filtered to include only biallelic Single Nucleotide Polymorphism (SNPs) with Variant Quality Score Log- Odds (VQSLOD)&#x2009;>&#x2009;1 and "PASS" status. Genetic diversity, differentiation, and selection signatures were analyzed using population genetics metrics. RESULTS: After filtering, 1312 samples were retained. Wright's inbreeding coefficient (Fws) showed that 875 (66.7%) samples had monoclonal infections, with the highest proportion of monoclonal infections in Ethiopia (95.3%), followed by Tanzania (67.2%), Kenya (65.7%), and Uganda (50%). Among the 875 monoclonal samples, 88 haplotypes were identified, with Hap_1 (renamed PF3D7)&#xa0;and Hap_13 comprising 37.9 and 21.5 of the samples, respectively. Nucleotide and haplotype diversity were relatively higher in Kenya with 0.097, and 0.88 respectively, compared to the other study populations. The overall fixation index (Fst) was&#x2009;<&#x2009;0.05, and Principal Component Analysis revealed no clear population sub-structure among countries. Negative Tajima's D values in Tanzania, Kenya, and Ethiopia indicated an excess of low-frequency alleles. CONCLUSION: This study reports low genetic diversity of the phistb gene in the four countries despite varying malaria transmission intensities among them, thus making it a suitable candidate gene for malaria vaccine. Further studies should be conducted to assess individual antibodies recognition of the phistb variants and the ability to elicit cross reactivity to further support its potential as a vaccine candidate.

Plasmodium falciparum

Mutational signature stratification of recurrent gliomas reveals distinct patterns of genomic traits.

BACKGROUND: Although temozolomide (TMZ) is widely used for glioma treatment, its therapeutic benefit is limited by acquired resistance and recurrence, facilitated by intratumor heterogeneity. Mutational signatures (MSs) inform tumor evolution and reveal alterations associated with treatment response. METHODS: We performed molecular analyses of 96 glioma recurrences with sufficient private single-nucleotide variants relative to their matched primary tumors, stratified by their dominant MS. RESULTS: Four groups were identified: MS11/TMZ-related (n&#x2009;=&#x2009;38), MS1/5/aging-related (n&#x2009;=&#x2009;32), MS6/15/21/26/microsatellite instability (MSI)-related (n&#x2009;=&#x2009;13), and other MS-related recurrences (n&#x2009;=&#x2009;13). MS11/TMZ-related recurrences showed higher acquired mutational counts than the other groups (1338 vs 59 (MS1/5/aging) vs 57 (MS6/15/21/26/MSI) vs 57 (other MSs); P&#x2009;<&#x2009;.01). Mutations in SYNE2, SZT2, and FBN3 were restricted to recurrences with dominant or second-dominant MS11/TMZ-related signature (n&#x2009;=&#x2009;41), and 85% (35/41) harbored mutations in these genes. In MS11/TMZ-related recurrences with RNA sequencing data (n&#x2009;=&#x2009;17), mRNA co-expression analyses identified SYNE2-ATAD5 and SZT2-MAPKBP1 associations. Among MS11/TMZ-related recurrences, MS23 was frequent (44%, 18/41) and associated with higher acquired mutational counts (2089 vs 1188; P&#x2009;=&#x2009;.018) and more IDH-wildtype tumors (67% vs 30%; P&#x2009;=&#x2009;.037). MAPKBP1 mutations were enriched in MS23-positive recurrences (56% (10/18) vs 0% (0/23); P&#x2009;<&#x2009;.001). MS1/5/aging-related recurrences showed more frequent acquired chromosome 16q losses (22% vs 8% (TMZ) vs 0% (MSI) vs 0% (other); P&#x2009;<&#x2009;.05), which were associated with an increased fraction of genome altered relative to 16q-diploid cases (15% vs 7%; P&#x2009;=&#x2009;.01). CONCLUSIONS: These findings show that MS-based stratification of recurrences refines molecular characterization after therapy and nominates candidate biomarkers and pathways for functional studies of treatment-associated glioma evolution.

mutational signatures

Effects of domestication on the body morphology and genetic diversity of the yellowfin seabream (Acanthopagrus latus).

The yellowfin seabream (Acanthopagrus latus) is a significant economic fish along the southeast coast of China. Recently, the drastic decline in the wild populations, exacerbated by overfishing and climate change, has heightened our reliance on aquaculture. However, the current lack of research on its domestication hinders effective conservation of wild populations and balanced management alongside the aquaculture industry. Studies on body characteristics have shown that wild yellowfin seabream possess a higher body, while cultured ones exhibit a wider body. Whole-genome SNP analysis revealed moderate genetic differentiation between cultured and wild populations. Further analyses of linkage disequilibrium, heterozygosity, and genetic diversity revealed that the degree of SNP linkage was lower in the wild population compared to the cultured population. In contrast, heterozygosity and nucleotide polymorphisms were significantly higher in the wild population (P&#xa0;<&#xa0;0.001 and P&#xa0;<&#xa0;0.05, respectively). Additionally, over 300 candidate genes were identified in each cultured population through genomic selection signature analysis, with 67 key genes shared among all three, which were linked to growth and development (ghrb, ghsra, and cfl1), immune response (aire, cd36, and igbp1), and salinity adaptation (abcc3, clic4, and kcnk15). Enrichment analysis indicated that the key candidate genes were significantly enriched in pathways related to protein kinase activity, ion binding and growth hormone synthesis, secretion and action (FDR&#xa0;<&#xa0;0.05). The findings provide valuable insights into the variation in body size of yellowfin seabream under domestication selection and offer an important theoretical basis for the genetic improvement of yellowfin seabream.

Animals

Transgene sequence codon optimization and composition determines replication competence of self-amplifying RNA.

Self-amplifying RNA (saRNA) is an emerging RNA therapeutic modality that can facilitate higher magnitude and more durable protein expression at substantially lower doses than nonreplicating mRNA. Unlike conventional messenger RNA (mRNA), alphavirus-derived saRNA must support a replicase-driven RNA amplification step in addition to translation, raising the possibility that transgene coding sequences impose sequence-level constraints on replication. Here, saRNA replication was found to be dependent on the codon composition of the transgene; multiple therapeutic transgenes were replication defective despite an intact Venezuelan Equine Encephalitis Virus (VEEV)-derived saRNA backbone. Replication defects were rescued by synonymous codon re-optimization of the same transgenes, indicating that nucleotide-level features of the coding sequence, rather than the encoded protein, govern replication competence. Comparative compositional analyses identified a distinct signature associated with productive replication, characterized by elevated GC (>53%) and GC3 (>63%) content, higher codon adaptation to human (>0.75), and reduced UpA (<43/kb) and UpU (<41/kb) dinucleotide density. Moreover, deliberate compositional perturbation of an otherwise replication-competent transgene shifted these features and abolished replication, supporting a causal and combinatorial role for sequence composition in defining saRNA replication outcome. These findings define an underappreciated constraint in saRNA therapeutics and motivate saRNA-specific payload design frameworks that incorporate alphavirus-associated compositional biases during transgene sequence optimization.

Codon

IMPACT OF FLUORESCENT DYES ON MUTATIONS IN NEXT GENERATION SEQUENCING LIBRARY GENERATION.

DNA labelling fluorescent dyes such as ethidium bromide have long been considered to be highly mutagenic during DNA replication. While recent studies have pushed back on this narrative, the intercalative nature of these dyes continues to raise the possibility that these dyes can induce mutations. The iconPCR instrument by n6tec uses fluorescent dyes to measure amplification in real time and to adjust cycling conditions. However, since this use of qPCR is preparative and not analytical, mutations introduced by fluorescent dyes would be propagated into the sequencing reaction. To address the impact of these dyes on downstream analyses, we have performed routine mutation calling as well as mutational signature analysis on samples amplified using the iconPCR in the presence of either SYBR or EvaGreen. Sequence analysis revealed very minimal impacts of dyes on the reactions, largely within the noise regimen with only subtle changes in mutation rates seen. Mutational signature analysis was unable to identify any key signatures assignable to the dyes in either substitutions or indel domains. The mutational impact of intercalating dyes during fluorescence-guided amplification is therefore minimal and can be disregarded in all but the most sensitive NGS applications.

Fluorescent Dyes