PubMed HealthSearch

SEARCH · PubMed Health

Results for “synonymous variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Hurdles to horizontal gene transfer: species-specific effects of synonymous variation and plasmid copy number determine antibiotic resistance phenotype.

Could codon composition condition the immediate success and the orientation of horizontal gene transfer? Horizontal gene transfer represents a change in the genome of expression of the transferred gene, and experimental evidence has accumulated indicating that the codon composition of a sequence is an important determinant of its compatibility with the translation machinery of the genome in which it is expressed. This suggests that codon composition influences the phenotype and the fitness conferred by a transferred gene and thus the immediate success of the transfer. To directly test this hypothesis, we characterized the resistance conferred by synonymous variants of a gentamicin resistance gene in three bacterial species: Escherichia coli, Acinetobacter baylyi and Pseudomonas aeruginosa. The strongest determinant of the resistance level conferred was the species in which the resistance gene was transferred, very likely because of important differences in the copy number of the plasmid carrying the gene. Significant differences in resistance were also found between synonymous variants within each of the three species, but more importantly, there was a strong interaction between species and variant: variants conferring high resistance in one species confer low resistance in another. However, the similarity in codon usage between the synonymous variants and the host genome only explained part of the phenotypic differences between variants in one species, P. aeruginosa. Further investigation of alternative explanations did not reveal common universal mechanisms across our three bacterial species. We conclude that codon composition can be a determinant of post-horizontal gene transfer success. However, there are multiple paths leading from synonymous sequence to phenotype, and sensitivity to these different paths is species-specific.

Gene Transfer, Horizontal

Analysis of human papillomavirus type 16 E4, E5 and L2 gene variations among women with cervical infection in Xinjiang, China.

BACKGROUND: There is a high incidence of cervical cancer in Xinjiang. Genetic variation in human papillomavirus may increase its ability to invade, spread, and escape host immune response. METHODS: HPV16 genome was sequenced for 90 positive samples of HPV16 infection. Sequences of the E4, E5 and L2 genes were analysed to reveal sequence variation of HPV16 in Xinjiang and the distribution of variation among the positive samples of HPV16 infection. RESULTS: Eighty-one of the 90 samples of HPV16 infection showed variation in HPV16 E4 gene with 18 nucleotide variation sites, of which 8 sites were synonymous variations and 11 missense variations. 90 samples of HPV16 infection showed variation in HPV16 E5 and L2 genes with 16 nucleotide variation sites (6 synonymous, 11 missense variations) in the E5 gene and 100 nucleotide variation sites in L2 gene (37 synonymous, 67 missense variations). The frequency of HPV16 L2 gene missense variations G3377A, G3599A, G3703A, and G3757A was higher in the case groups than in the control groups. CONCLUSIONS: Phylogenetic tree analysis showed that 87 samples were European strains, 3 cases were Asian strains, there were no other variations, and G4181A was related to Asian strains. HPV16 L2 gene missense variations G3377A, G3599A, G3703A, and G3757A were significantly more frequent in the case groups than in the control groups.

Humans

FLT4 gene polymorphisms influence isolated ventricular septal defect predisposition in a Southwest China population.

BACKGROUND: Ventricular septal defect (VSD) is the most common congenital heart disease. Although a small number of genes associated with VSD have been found, the genetic factors of VSD remain unclear. In this study, we evaluated the association of 10 candidate single nucleotide polymorphisms (SNPs) with isolated VSD in a population from Southwest China. METHODS: Based on the results of 34 congenital heart disease whole-exome sequencing and 1000 Genomes databases, 10 candidate SNPs were selected. A total of 618 samples were collected from the population of Southwest China, including 285 VSD samples and 333 normal samples. Ten SNPs in the case group and the control group were identified by SNaPshot genotyping. The chi-square (&#x3c7;2) test was used to evaluate the relationship between VSD and each candidate SNP. The SNPs that had significant P value in the initial stage were further analysed using linkage disequilibrium, and haplotypes were assessed in 34 congenital heart disease whole-exome sequencing samples using Haploview software. The bins of SNPs that were in very strong linkage disequilibrium were further used to predict haplotypes by Arlequin software. ViennaRNA v2.5.1 predicted the haplotype mRNA secondary structure. We evaluated the correlation between mRNA secondary structure changes and ventricular septal defects. RESULTS: The &#x3c7;2 results showed that the allele frequency of FLT4 rs383985 (P&#x2009;=&#x2009;0.040) was different between the control group and the case group (P&#x2009;<&#x2009;0.05). FLT4 rs3736061 (r2&#x2009;=&#x2009;1), rs3736062 (r2&#x2009;=&#x2009;0.84), rs3736063 (r2&#x2009;=&#x2009;0.84) and FLT4 rs383985 were in high linkage disequilibrium (r2&#x2009;>&#x2009;0.8). Among them, rs3736061 and rs3736062 SNPs in the FLT4 gene led to synonymous variations of amino acids, but predicting the secondary structure of mRNA might change the secondary structure of mRNA and reduce the free energy. CONCLUSIONS: These findings suggest a possible molecular pathogenesis associated with isolated VSD, which warrants investigation in future studies.

Child

Genetic evidence for predisposition to acute leukemias due to a missense mutation (p.Ser518Arg) in ZAP70 kinase: a case-control study.

BACKGROUND: The apparent lack of additional missense mutations data on mixed-phenotype leukemia is noteworthy. Single amino acid substitution by these non-synonymous single nucleotide variations can be related to many pathological conditions and may influence susceptibility to disease. This case-control study aimed to unravel whether the ZAP70 missense variant (rs104893674 (C&#x2009;>&#x2009;A)) underpinning mixed-phenotype leukemia. METHODS: The rs104893674 was genotyped in clients who were mixed-phenotype acute leukemia-, acute lymphoblastic leukemia- and acute myeloid leukemia-positive and matched healthy controls, which have been referred to all major urban hospitals from multiple provinces of country- wide, IRAN, from February 11' 2019 to June 10' 2023, by amplification refractory mutation system-polymerase chain reaction method. Direct sequencing for rs104893674 of the ZAP70 gene was performed in a 3130 Genetic Analyzer. RESULTS: We found that the AC genotype of individuals with A allele at this polymorphic site (heterozygous variant-type) contribute to the genetic susceptibility to acute leukemia of both forms, acute myeloid leukemia and acute lymphoblastic leukemia as well as with a mixed phenotype. In other words, the ZAP70 missense variant (rs104893674 (C&#x2009;>&#x2009;A)) increases susceptibility of distinct cell populations of different (myeloid and lymphoid) lineages to exhibiting cancer phenotype. The results were all consistent with genotype data obtained using a direct DNA sequencing technique. CONCLUSION: Of special interest are pathogenic missense mutations, since they generate variants that cause specific molecular phenotypes through protein destabilization. Overall, we discovered that the rs104893674 (C&#x2009;>&#x2009;A) variant chance in causing mixed-phenotype leukemia is relatively high.

Humans

Population history rather than tree age contributes to the evolutionary importance of ancient trees in an endangered conifer.

Ancient trees are in global decline and face increasing conservation challenges. Their exceptional longevity has fostered the view that they are genetic reservoirs, yet whether old age is synonymous with unique genetic variation remains unclear. Here we assembled a ~8-Gb chromosome-level reference genome for the critically endangered conifer Glyptostrobus pensilis, now largely restricted to southern China with scattered populations in Vietnam and Laos, and resequenced 147 individuals, including 64 ancient (>100&#x2009;years old and persisting in human-dominated landscapes), 33 wild and 50 recently cultivated individuals. Ancient individuals comprised both likely natural relics and historically introduced individuals and formed two deeply divergent lineages and one ancestral-admixed group, each with distinct demographic histories of prolonged contraction and genomic erosion. Lineage identity explained more variation in genome-wide diversity, inbreeding and genetic load than the three conservation types, despite broad differences in age structure. Rare-allele analyses revealed pronounced heterogeneity among ancient trees: only relic and ancestral-origin individuals from high-diversity lineages contributed substantial unique variation, much of which is poorly represented in wild and cultivated populations. Together, our findings suggest that ancient trees are not uniformly genetically irreplaceable and that, at least in this conifer, evolutionary importance is shaped more strongly by population history than by age alone.

Endangered Species

Human T cell epitopes of Mycobacterium tuberculosis are evolutionarily hyperconserved.

Mycobacterium tuberculosis is an obligate human pathogen capable of persisting in individual hosts for decades. We sequenced the genomes of 21 strains representative of the global diversity and six major lineages of the M. tuberculosis complex (MTBC) at 40- to 90-fold coverage using Illumina next-generation DNA sequencing. We constructed a genome-wide phylogeny based on these genome sequences. Comparative analyses of the sequences showed, as expected, that essential genes in MTBC were more evolutionarily conserved than nonessential genes. Notably, however, most of the 491 experimentally confirmed human T cell epitopes showed little sequence variation and had a lower ratio of nonsynonymous to synonymous changes than seen in essential and nonessential genes. We confirmed these findings in an additional data set consisting of 16 antigens in 99 MTBC strains. These findings are consistent with strong purifying selection acting on these epitopes, implying that MTBC might benefit from recognition by human T cells.

Antigens, Bacterial

Evolutionary dynamics of the chloroplast genome in Abutilon (Malvoideae, Malvaceae).

The genus Abutilon Mill. (Malvaceae) comprises approximately 178 species distributed across tropical and subtropical regions, many of which hold significant ornamental, economic, and medicinal value; yet its taxonomic classification remains challenging. In this study, six species were sequenced from herbarium specimens, and the chloroplast (cp.) genomes of ten additional species were assembled de novo from publicly available raw data. Three previously reported cp. genomes were also incorporated to characterise cp. genome structure, identify polymorphic loci, and perform phylogenetic analyses. The cp. genomes ranged from 159,458 to 160,454&#xa0;bp and exhibited the typical quadripartite structure, with each genome containing 112 unique genes (78 protein-coding, 30 tRNA, and 4 rRNA) that showed conserved content and organisation. These genomes exhibited high similarity in GC content, inverted repeat boundaries, relative synonymous codon usage, amino acid frequencies, and substitution patterns. However, notable variation was observed in the total number of simple sequence repeats, ranging from 70 to 97 per genome. Selection analyses indicated predominant purifying selection, with evidence of episodic positive selection detected in rpoC2, rbcL, and ycf1. Two codons in rbcL were clade-specific and provided phylogenetic signal distinguishing Australian and Old World pantropical species. Nucleotide diversity analysis identified six highly polymorphic intergenic spacers (trnH-psbA, rps19-rpl2, psbT-pbf1, psaC-ndhD, trnR-atpA, and ndhJ-ndhK) that may be suitable for taxonomic studies. The phylogeny from maximum likelihood (ML) and Bayesian inference (BI) resolved two major clades: one comprising an exclusively Australian lineage occurring predominantly in arid and semi-arid environments, and the other a pantropical lineage spanning multiple continents. Abutilon grandifolium was recovered as sister to the remaining sampled Abutilon taxa in both ML and BI analyses, although no biogeographic origin inference can be drawn from this placement pending broader taxon sampling and integration of nuclear genomic data. These findings provide insights into the evolutionary dynamics of the cp. genome in Abutilon and offer a foundational genomic framework for refining Abutilon taxonomy.

Genome, Chloroplast

Comparative Genomic Analysis of Six Mycoplasma Gallisepticum Strains: Insights into Genetic Diversity and Antibiotic Resistance.

Mycoplasma gallisepticum (MG) is a significant pathogen that causes respiratory diseases, which have had a substantial economic impact on the poultry industry. Despite the resistance of MG to antibiotics, it is imperative to identify genetic diversity in order to develop countermeasures. In this study, the genomes of six MG strains were examined to gain deeper insights into the mutations. The data pertaining to Variant Annotation and Mutation Analysis using SnpEff, along with the calculation of mutation rates as the ratio of total mutations to the length of the genomic regions analyzed, were thoroughly examined. The comprehensive evaluation yielded a total of 25,942 variants across the six strains, underscoring substantial genetic diversity. Notably, strain S6 exhibited a preponderance of frameshift mutations. A notable finding was the presence of a mutation in the MsbA gene shared by all six strains. Furthermore, five of the six strains, with the exception of strain F99 Lab, exhibited a mutation at position 5158, which impacts a multidrug transport system. Notably, strain ATCC exhibits a distinctive mutation at position 942, while strain S6 displays a unique mutation at position 6855, which is linked to efflux ABC transporter components. Furthermore, a substantial degree of genetic variation was observed among the CrmA, GapA, and vlhA genes among the various strains. High-impact changes, such as insertions and deletions, exhibited a higher frequency in CrmA, particularly in strain S6. Conversely, nonsynonymous variations demonstrated a heightened prevalence in GapA, particularly in strain F99 Lab. The vlhA gene exhibited a spectrum of effects, ranging from synonymous mutations to high-impact mutations such as stop-gains and frameshifts, particularly in strains k5111a and k4602. The functional variations observed among the strains can be attributed to these mutations, which have the potential to alter gene expression or protein function. Furthermore, substantial mutations in the dxr and rpoC genes were associated with antibiotic resistance. These mutations underscore the ongoing evolutionary adaptations of M. gallisepticum. Consequently, there is an imperative for the revision of treatment protocols and the formulation of targeted vaccines to regulate resistance within the poultry industry.

Mycoplasma gallisepticum

Rare pathogenic NR2F2 (COUP-TFII) variants as potential etiological causes in pediatric patients with congenital heart diseases (CHDs).

OBJECTIVES: Congenital heart diseases (CHDs) are complex genetic disorders, and their genetic basis is not yet fully understood. Nuclear receptor subfamily 2 group F member 2 (NR2F2 or COUP-TFII) encodes a transcription factor which is expressed at high levels during mammalian development. Few studies have identified heterozygous and rare variants in the NR2F2 gene in individuals with CHD. This study aimed to evaluate the association between pathogenic genetic alterations in NR2F2 with CHD risk. METHODS: A case-control study was conducted on a group of 135 patients (83 boys and 52 girls) with various types of non-hereditary, isolated CHD who were undergoing open-heart surgery. Additionally, 95 matched healthy children without syndromic or isolated heart abnormalities were selected. RESULTS: Using Sanger sequencing, we identified 5 heterozygous single nucleotide variants in exons 2 and 3 of the NR2F2 gene. These variations were novel and not present in any genomic variation databases. Four of the variations were missense mutations (p.Pro159Arg, p.Ser329Phe, p.Qln338Pro, and p.Tyr348Ser) and one was a synonymous variant (p.G361 = ) in the coding region. Importantly, in silico results indicated that the missense variants had pathogenic effects on protein function. Additionally, the missense variants substantially altered the predicted structure of COUP-TFII. CONCLUSION: The results we obtained not only validate the correlation between NR2F2 mutations and CHDs but also have significant potential for guiding new preventive and therapeutic strategies. This could contribute to the advancement of medical interventions in the fields of cardiology and genetics.

Humans

Selection profiles in RNA viruses reflect the characteristics of viruses more than individual proteins.

Proteins that are exposed on the surface of a virus are frequently subject to strong selection to escape from neutralizing antibodies. To investigate whether surface-exposed (SE) and non-exposed (NE) proteins encoded by RNA viruses exhibit different patterns of evolution under selection, we analyzed 244 protein-coding genes from 28 species of RNA viruses representing 15 taxonomic families. First, we show that gene-wide rates of non-synonymous (dN) and synonymous (dS) substitutions do not differentiate between SE and NE proteins. To incorporate variation in substitution rates among codon sites, we inferred the posterior distribution over a fixed grid of dN and dS rates for each alignment. This 'evolutionary fingerprint' provides a common framework for comparing the selection profiles of non-homologous genes. Next, we computed the Wasserstein distance for every pair of fingerprints, which is analogous to amount of work required to reshape one distribution to another. After compensating for differences in genetic variation among alignments, we found a small but significant difference between the fingerprints of SE and NE proteins (PERMANOVA, P&#x2009;=&#x2009;0.03). However, we observed larger and more significant effects of whether the virus is enveloped (P&#x2009;<&#x2009;10-5) and the interaction between these factors (P=6.9&#xd7;10-4). The latter effects were driven by high levels of purifying selection in capsid proteins of Picornaviruses. Furthermore, greater amounts of variation in fingerprints were explained by significant differences among virus families and modes of transmission (P&#x2009;<&#x2009;10-5). These results imply the pattern of selection on a virus protein is shaped more by characteristics of the virus than the protein itself.

RNA Viruses

Comprehensive plastome variation and RNA editing in Mentha: insights into phylogenetic relationships and candidate DNA barcodes.

INTRODUCTION: Mentha is an economically and medicinally important genus in Lamiaceae, but its taxonomy and species delimitation remain challenging because of frequent hybridization, polyploidy, and marked morphological plasticity. METHODS: In this study, we comparatively analyzed 12 plastomes representing major Mentha species, hybrid taxa, and unresolved accessions, including four newly assembled genomes, to characterize plastome structure, repeat composition, sequence divergence, phylogenetic relationships, and plastid RNA editing. The M. arvensis plastome and RNA-seq datasets originated from independent Swiss and Indian accessions, respectively. RESULTS: The plastomes were highly conserved in overall organization, ranging from 151,824 to 152,154 bp and displaying the typical quadripartite structure. Gene content and order were largely stable across taxa, with only minor variation likely associated with annotation differences at IR/SC boundary regions. Codon usage analysis revealed a clear bias toward A/U-ending synonymous codons, and most shared protein-coding genes showed low Ka/Ks ratios, indicating predominant purifying selection. Repeat analyses showed that simple sequence repeats were mainly composed of A/T-rich mononucleotide motifs, whereas long repeats were concentrated in the 30-40 bp size class. Comparative analyses identified six hypervariable regions, namely ccsA-ndhD, ycf1, ndhD, rpl32-trnL-UAG, rbcL-accD, and petA-psbJ, which represent promising candidate plastid markers for species discrimination. Phylogenetic analysis based on complete plastomes provided strong support for relationships among the sampled taxa and recovered a close affinity among M. aquatica, M. arvensis, and M. canadensis. In addition, RNA-seq analysis of M. arvensis identified 17 candidate plastid RNA editing sites, most of which were C-to-U conversions and nonsynonymous events. DISCUSSION: Together, these results expand plastid genomic resources for Mentha and provide a useful framework for phylogenetic inference, species identification, and future germplasm utilization.

RNA editing

Codon bias variation in Staphylococcus aureus.

BACKGROUND: Staphylococcus aureus causes a multiplicity of human diseases acquired in community and healthcare settings alike around the globe. While most studies focus on coding changes to assess genome evolution and study genetic adaptation, interrogation of silent mutations in the form of synonymous codon usage bias is less well-studied. As such, understanding of patterns in codon bias at the gene and genome levels, and how codon bias impacts protein expression in S. aureus remains incomplete. METHODS: The codon bias of 2,565 protein encoding genes from NCTC 8325 was queried against all publicly available closed S. aureus genomes. Using public BioSample data, genomes were sorted by disease state, submitting institution, and collection site. Codon bias was assessed at the level of gene and genome using the codon adaptation index (CAI), calculated using 30S and 50S ribosomal genes. Gene set enrichment analysis was applied to determine associations between physiological functions, CAI gene scores, and interquartile ranges. CAI scores were also compared to an in vitro S. aureus proteomics database to correlate codon bias and protein expression. RESULTS: CAI scores varied within and between isolates at the gene and genome levels. Genes with ribosome-associated functions were most enriched among high CAI genes, and had low CAI interquartile ranges (IQR), suggesting selective pressure to maintain high expression of these genes across all S. aureus isolates. Genome sequences submitted by Aga Khan University Hospital, Nairobi, Kenya were most different from others. For the LAC USA 300 strain, CAI and protein expression were moderately positively correlated (cor&#x2009;=&#x2009;0.534, p&#x2009;<&#x2009;2.2e-16). CONCLUSIONS: Codon bias in S. aureus was shown to vary between gene, and to be a source of genetic variation between isolates; CAI and in vitro protein expression were positively correlated.

Staphylococcus aureus

Mitogenome assembly and phylogenetic relationships of Phalaris arundinacea.

INTRODUCTION: As a perennial herb of Poaceae, Phalaris arundinacea plays key roles in grazing, production, and soil and water conservation because of its well-developed rhizomes and seed dispersal. We assembled and annotated the first mitogenome of P. arundinacea to support evolutionary and taxonomic research. METHODS: We assembled and annotated the first complete mitochondrial genome of P. arundinacea by integrating Illumina short reads with Nanopore long reads via a hybrid assembly strategy. The genome architecture was comprehensively characterized, encompassing codon usage bias, repetitive sequence organization, and inter-organellar genetic exchange with the chloroplast genome. RESULTS AND DISCUSSION: Assembly of the P. arundinacea mitogenome revealed two circular structures with a combined length of 526,717 bp. The genome comprised a set of 37 protein-coding genes (PCGs), 27 tRNAs, and 8 rRNAs, with the rRNA genes exhibiting full assembly (100% coverage). The mitochondrial genome contained 154 forward and 164 palindromic repeats, along with 25 tandem repeats and 124 simple sequence repeats (SSRs). Notably, 102 SSRs were distributed on contig1, predominantly in tetrameric form. Furthermore, 376 RNA editing sites were predicted. A total of 104 fragments were integrated into the mitochondrial genome from the chloroplast, amounting to 55,866 bp of transferred sequence. Finally, phylogenetic analysis of 28 plant mitogenomes placed P. arundinacea closest to species within the genus Poa (P. chaixii and P. pratensis). Comparative analysis of non-synonymous-to-synonymous substitution rate (Ka/Ks) ratios across divergent species revealed that the mitochondrial genome of P. arundinacea underwent stabilizing evolutionary dynamics, characterized by predominant purifying selection with several lineage-specific variations in selective pressure. Our findings support the close phylogenetic relationship between P. arundinacea and species of the genus Poa and provide a reference mitochondrial genome resource for future comparative studies within Phalaris that incorporate broader taxon sampling. These results support deeper phylogenetic investigations of P. arundinacea and facilitate future work on its germplasm characterization and applied use.

Phalaris arundinacea

Genome-wide variation analysis of two Salvia hispanica L. genotypes and implication for associations with metabolic and adaptive traits.

BACKGROUND: Advances in next-generation sequencing have accelerated genome-wide exploration of genetic diversity in underutilized oilseed crops. Salvia hispanica L. (chia), a high-nutrient pseudocereal rich in omega-3 fatty acids, is increasingly valued for its health benefits and commercial potential, yet it remains poorly characterized at the genomic level. Understanding the scale and nature of genomic variation is essential for improving complex traits such as oil yield, stress tolerance, and seed quality. METHODS: Two contrasting chia genotypes, Black-chia (CACH-B) and White- chia (CACH-W), were resequenced using the Bio-Resequencing Toolkit (BRT) pipeline. High-coverage sequencing, with a mapping rate exceeding 99% and an average depth of approximately 28&#xd7;, facilitated the detection and annotation of single-nucleotide polymorphisms (SNPs), insertions and deletions (InDels), copy-number variations (CNVs), and structural variants (SVs). The functional classification of variant impacts enabled the identification of genes potentially linked to metabolic and adaptive traits. RESULTS: A total of 1.97 million SNPs, 401,493 InDels, 836 CNVs, and 15,288 SVs were identified across the chia genome. Notably, approximately 53% of exonic SNPs were non-synonymous (dN/dS&#xa0;&#x2248;&#xa0;1.28), predominantly affecting lipid metabolism, transcriptional regulation, and stress response pathways, potentially altering key agronomic traits. In addition, CNV hotspots were concentrated in chromosomes 3 and 6, overlapping MYB, WRKY, and bZIP transcription factor loci, may potentially be involved in stress tolerance and yield. Furthermore, structural rearrangements, including inversions and duplications within the FAD2, FAD3, and CYP450 gene clusters, were potentially associated with seed pigmentation and omega-3 biosynthesis, pointing to their potential breeding relevance. Observed heterozygosity (H&#x2092;&#xa0;&#x2248;&#xa0;0.71) and nucleotide diversity (&#x3c0;&#xa0;&#x2248;&#xa0;7&#xa0;&#xd7;&#xa0;10-3) indicated moderate to high allelic richness. In addition, the low FST value (0.038) indicates substantial genomic similarity between the two genotypes. CONCLUSION: This study presents the first comprehensive map integrating SNPs, CNVs, and SVs in S. hispanica L. The results reveal a structurally dynamic genome characterized by substantial sequence and structural variation, providing valuable insights into genomic diversity and potential adaptive mechanisms in chia. The coexistence of high SNP diversity and abundant structural variation underpins chia's nutritional specialization and environmental resilience. These results deliver a foundational genomic resource for marker-assisted breeding, genome-wide association studies, and the development of climate-resilient chia cultivars.

Copy-number variation, structural variation

Exploring the Mitochondrial Genomes of Phoebe Species (Lauraceae): Structural Dynamics and Functional Conservation.

Plant mitochondrial genomes (mitogenomes) vary markedly in size and architecture despite generally slow rates of sequence evolution. Phoebe is an ecologically and economically valuable genus of Lauraceae, yet its mitogenome diversity remains poorly characterized. In this study, we newly sequenced, assembled, and annotated the mitogenomes of three nationally protected Class II wild plants (P. bournei, P. chekiangensis, P. zhennan) from China and compared their mitogenomic characteristics. The three assemblies were resolved into representative circular configurations ranging from 808 to 864&#x2009;kb, with similar GC contents and conserved protein-coding capacity. Each mitogenome contained distinct 41 protein-coding genes, 27-28 transfer RNAs, and three ribosomal RNAs. Synteny analysis revealed extensive changes in homologous-block order and orientation despite substantial sequence homology among the three species. Abundant repeats occurred predominantly in noncoding regions, while plastid-derived fragments documented historical intracellular DNA transfer. The three species exhibited similar codon usage and predicted RNA-editing patterns, whereas low synonymous divergence limited inference from pairwise ratios. Phylogenetic analysis based on mitochondrial protein-coding genes recovered Phoebe as a well-supported monophyletic lineage. These results reveal substantial structural divergence accompanied by conserved nucleotide composition and coding capacity, providing valuable data for further understanding the evolutionary variation of plant mitogenomes of Phoebe and the Lauraceae.

Phoebe

Compensatory Evolution Following Deleterious Episodes of GC-biased Gene Conversion in Rodents.

GC-biased gene conversion (gBGC) is a widespread evolutionary force associated with meiotic recombination that favors the accumulation of deleterious AT to GC substitutions in proteins, moving them away from their fitness optimum. In many mammals, recombination hotspots have a rapid turnover, leading to episodic gBGC, with the accumulation of deleterious mutations stopping when the recombination hotspot dies. Selection is therefore expected to act to repair the damage caused by gBGC episodes through compensatory evolution. However, this process has never been studied or quantified so far. Here, we analyzed the nucleotide substitution pattern in coding sequences of a highly diversified group of Murinae rodents. Using phylogenetic analyses of about 70,000 coding exons, we identified numerous exon-specific, lineage-specific gBGC episodes, characterized by a clustering of synonymous AT to GC substitutions and by an increasing rate of nonsynonymous AT to GC substitutions, many of which are potentially deleterious. Analyzing the molecular evolution of the affected exons in downstream lineages, we found evidence for pervasive compensatory evolution after deleterious gBGC episodes. Compensation appears to occur rapidly after the end of the episode and to be driven by the standing genetic variation rather than new mutations. Our results demonstrate the impact of gBGC on the evolution of amino-acid sequences and underline the key role of epistasis in protein adaptation. This study contributes to a growing body of literature emphasizing that adaptive mutations, which arise in response to environmental changes, are just 1 subset of beneficial mutations, alongside mutations resulting from oscillations around the fitness optimum.

Gene Conversion

Epstein-Barr Virus Sequence Variations Among the Understudied Nasopharyngeal Carcinoma Patients of Diverse Ancestries in Southeast Asia.

Epstein-Barr virus (EBV) is associated with cancers, including lymphomas and nasopharyngeal carcinoma (NPC). To date, risk variants for NPC were mainly identified from Chinese populations, which dominated the world's total number of cases. Although Southeast Asia (SEA) countries have among the world's top yet intriguingly diverse NPC age-standardized incidence rates across subpopulations, data on EBV from SEA remains scarce. In this study, we examined 83 NPC patients of different ancestries for the presence of risk haplotypes associated with the Southern Chinese NPC and generated and analyzed 67 EBV sequences (from tissue, patient-derived xenografts and lymphoblastoid cell lines of 60 NPC patients) together with 838 published EBV genomes. Our study revealed that NPC patients of non-Chinese ancestry had fewer risk variants and haplotypes that are associated with Southern Chinese NPC and clustered distinctly from lymphomas, Southern Chinese NPC, and non-cancer controls. The distribution of non-synonymous variants was similar among NPC patients of Chinese ancestry, irrespective of geographical location. Meanwhile, non-synonymous variants in genes related to packaging, latency, and structural proteins such as BPLF1, LF3, and LMP1 varied across different ancestries. Our findings suggest possibilities of EBV adaptation to host genetics for NPC pathogenesis and warrant further research for the understudied NPC subpopulations.

Adult

The human IG heavy chain constant gene locus is enriched for large structural variants and coding polymorphisms that vary among human populations.

The human immunoglobulin heavy chain constant (IGHC) domain of antibodies (Ab) is responsible for effector functions critical to immunity. This domain is encoded by genes in the IGHC locus, where descriptions of genomic diversity remain incomplete. We utilized long-read sequencing to build an IGHC haplotype/variant catalog from 105 individuals of diverse ancestry. We discovered uncharacterized single nucleotide variants (SNV) and large structural variants (SVs, n=7), representing new genes and alleles enriched for non-synonymous substitutions, highlighting potential functional effects. Of the 221 identified IGHC alleles, 192 were novel. SNV, SV, and gene allele/genotype frequencies revealed population differentiation, including (i) hundreds of SNVs in African and East Asian populations exceeding a fixation index (FST) of 0.3, and (ii) an IGHG4 haplotype carrying coding variants uniquely enriched in Asian populations. Our results illuminate missing signatures of IGHC diversity and establish a new foundation for investigating IGHC germline variation in Ab function and disease.

Journal Article