PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genomic Structural Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Comparative analysis of chloroplast genomes in ten holly (Ilex) species: insights into phylogenetics and genome evolution.

In order to clarify the chloroplast genomes and structural features of ten Ilex species and provide insights into the phylogeny and genome evolution of the genus Ilex, we conducted a comparative analysis of chloroplast genomes using bioinformatics methods. The chloroplast genomes of ten Ilex species were obtained, and their structural features and variations were compared. The results indicated that all chloroplast genomes in the genus Ilex exhibit a double-stranded circular structure, with sizes ranging from 157,356 to 158,018 bp, showing minimal differences in size. The chloroplast genomes of the ten Ilex species have a relatively conservative gene count, with a total of 134 to 135 genes, including 88 or 89 protein-coding genes, and a conserved number of 8 rRNA genes. Each chloroplast genome contains 3 to 123 SSR (Simple Sequence Repeat) sites, predominantly composed of mononucleotide and trinucleotide repeats, with no detection of pentanucleotide or hexanucleotide repeats. The variation in dispersed repeat sequences among Ilex species is minimal, with a total repeat sequence number ranging from 1 to 14, concentrated in the length range of 30 to 42 base pairs. The expansion and contraction of chloroplast genome boundaries among Ilex species are relatively stable, with only minor variations observed in individual species. Variations in non-coding regions are more pronounced than those in coding regions, with the variability in the Large Single Copy region (LSC) being the highest, while the variability in the Inverted Repeat region A (IRa) is the lowest. The divergence time among Ilex species was estimated using the MCMC-tree module, revealing the evolutionary relationships among these species, their common ancestors, and their differentiation throughout the evolutionary process. The research findings provide a valuable reference for the systematic study and molecular marker development of Ilex plants.

Genome, Chloroplast

A pangenome framework uncovers the role of deletions in repeated evolution of cave-derived traits.

Structural variants (SVs) are increasingly recognized as key contributors to adaptive evolution, yet they remain underexplored compared with single-nucleotide variation. To understand how large-scale genomic changes shape repeated evolution, we leveraged multiple levels of sequence data across the powerful evolutionary model system of the Mexican tetra fish (Astyanax mexicanus). We constructed one of the first pangenome graphs from a naturally evolving vertebrate, enabling comprehensive discovery of SVs among 120 fish from 11 populations. We discover substantial amounts of structural variation and explore the roles of genomic biases and selection in shaping the distribution of these variants. More than 2400 high-confidence cave-specific deletions are enriched in biological pathways involved in vision, metabolism, and behavior and cluster nonrandomly in quantitative trait loci linked to cavefish traits. Additionally, 67 genes harbor unique deletions between independent cavefish lineages. These reused genes show evidence of population-specific selection (99% contain selective sweeps compared with 8%-15% in genes lacking SVs), indicating that deletions likely rose in frequency through repeated positive selection rather than drift. Together, these results reveal that recurrent deletion events have repeatedly contributed to the evolution of cave-adapted phenotypes and highlight deletions as underexplored contributors of adaptive evolution in extreme environments.

Animals

An allelic resolution gene atlas for tetraploid potato provides insights into tuberization and stress resilience.

Tubers are modified underground stems that enable asexual, clonal reproduction and serve as a mechanism for overwintering and avoidance of herbivory. Tubers are wide-spread across angiosperms with some species such as Solanum tuberosum L. (potato) serving as a vital crop for human consumption. Genes responsible for tuber initiation and disease resistance have been characterized in potato including StSP6A, a homolog of Flowering Time, that functions as tuberigen, the equivalent of florigen. To elucidate additional molecular and genetic mechanisms underlying potato biology including tuber initiation, tuber development, and stress responses, we generated a developmental and abiotic/biotic-stress gene expression atlas from 34 tissues and treatments of Atlantic, a tetraploid cultivar. Using the haplotype-phased tetraploid Atlantic genome assembly and expression abundances of 129,218 genes, we constructed gene coexpression modules that represent networks associated with distinct developmental stages as well as stress responses. Functional annotations were given to modules and used to identify genes involved in tuberization and stress resilience. Structural variation from a pan-genomic analysis across four cultivated potato genome assemblies as well as domestication and wild introgression data allowed for deeper insights into the modules to identify key genes involved in tuberization and stress responses. This study underscores the importance of transcriptional regulation in tuberization and provides a comprehensive framework for future research on potato development and improvement.

Journal Article

An allelic resolution gene atlas for tetraploid potato provides insights into tuberization and stress resilience.

Tubers are modified underground stems that enable asexual, clonal reproduction and serve as a mechanism for overwintering and avoidance of herbivory. Potato (Solanum tuberosum L.) is cultivated for its tubers, which serve as a major crop. Genes responsible for tuber initiation and disease resistance have been characterized in potato including StSP6A, a homolog of flowering time, that functions as a tuberigen, the equivalent of a florigen. To elucidate additional molecular and genetic mechanisms underlying potato biology including tuber initiation, tuber development, and stress responses, we generated a developmental and abiotic/biotic-stress gene expression atlas from 34 tissues and treatments of the tetraploid potato cultivar, Atlantic. Using the haplotype-phased tetraploid Atlantic genome assembly and expression abundances of 129 218 genes, we constructed gene coexpression modules that represent networks associated with distinct developmental stages as well as stress responses. Functional annotations were given to modules and used to identify genes involved in tuberization and stress resilience. Structural variation from a pan-genomic analysis across four cultivated potato genome assemblies as well as domestication and wild introgression data allowed for deeper insights into the modules to identify key genes involved in tuberization and stress responses. This study underscores the importance of transcriptional regulation in tuberization and provides a comprehensive framework for future research on potato development and improvement.

Solanum tuberosum

Genomic characterization of KPC-2 and NDM coproducing carbapenem-resistant Klebsiella pneumoniae in a hospital: discovery of ST1869 clone and a novel hybrid plasmid.

UNLABELLED: To characterize the plasmid architecture and molecular background of KPC-NDM coproducing carbapenem-resistant Klebsiella pneumoniae (KN-CRKP) in a South China hospital. Five KN-CRKP isolates were collected, including three from one patient. All underwent Illumina sequencing; two (ST11 and ST1869) additionally had Nanopore sequencing. Antimicrobial susceptibility testing strain sequence types, conjugation assays, resistance gene profiling, plasmid typing, genetic structure comparison, core-genome single nucleotide polymorphisms (SNPs) analysis, and plasmid clustering were performed. All isolates exhibited an imipenem minimum inhibitory concentration (MIC) of ≥128 µg/mL and harbored multiple resistance genes. One isolate (1/5) belonged to ST1869 and co-harbored blaKPC-2 and blaNDM-5. The blaNDM-5-carrying plasmid was a novel IncI1/X3 fusion plasmid that also carried blaCMY-42. Unlike several IncX3 plasmids carrying blaNDM in publicly available KN-CRKP genomes from South China, this IncI1/X3 hybrid lacked a complete conjugative transfer system. ST11 was the predominant clone (4/5), co-harboring blaKPC-2 and blaNDM-1. A rare genetic structure, ΔISKpn6-blaKPC-2-ISKpn28, was identified on IncFII plasmids carrying blaKPC-2. Plasmid clustering analysis of 126 comparative KN-CRKP genomes showed diverse sequence types and plasmid backgrounds associated with the KPC/NDM co-production pattern. The observed plasmid diversity and structural variation in KN-CRKP support continued genomic surveillance, with particular attention to the ST1869 clone, the novel IncI1/X3 hybrid plasmid harboring blaNDM-5 and blaCMY-42, and the rare "ΔISKpn6-blaKPC-2-ISKpn28" genetic structure. Expanded genomic data on KN-CRKP are needed to further elucidate its resistance mechanisms and plasmid evolutionary trajectories. IMPORTANCE: The co-production of KPC and NDM carbapenemases in Klebsiella pneumoniae poses a formidable threat to clinical antimicrobial therapy, as these enzymes confer resistance to virtually all β-lactam agents, including carbapenems. Here, we report novel genomic features of KN-CRKP in South China, including the emergence of the ST1869 clone, a unique IncI1/X3 hybrid plasmid harboring blaNDM-5 and blaCMY-42, and the rare ΔISKpn6-blaKPC-2-ISKpn28 genetic structure. These findings substantially expand current understanding of plasmid evolution and resistance gene dissemination in this region. The identification of diverse resistance mechanisms and clonal backgrounds supports enhanced genomic surveillance and infection-control awareness for pan-resistant Enterobacterales.

Plasmids

The Rise of Plant Pan-Genomes: From Genome Variation to Predictive Breeding.

Plant pan-genomics is entering a new phase beyond genome variation discovery, requiring a shift from cataloguing genomic diversity toward understanding how variation generates biological function and breeding value. Here, we propose that the future of plant pan-genomics will be shaped by three conceptual transitions. First, structural variation (SV), presence-absence variation (PAV), and haplotype diversity should be interpreted not merely as genomic differences, but as regulatory components that influence gene networks, chromatin organization, and complex traits. Second, the expansion from species-level pan-genomes to genus-level super pan-genomes provides an evolutionary framework for uncovering adaptive genetic modules preserved in wild relatives and overlooked during domestication. Third, integrating pan-genomes with pan-omics, three-dimensional genome analyses, and artificial intelligence will enable the transformation of genomic variation into predictive models for crop improvement. We further propose that the ultimate value of pan-genomes lies not in generating increasingly complete genome collections, but in establishing a mechanistic bridge between genome diversity, biological function, and breeding decisions. This transition will move crop improvement from empirical selection toward rational genome design, where evolutionary diversity can be systematically interpreted, predicted, and engineered.

Journal Article

Gene conversion confined to a direct repeat of the acceptor splice site generates allelic diversity at human glycophorin (GYP) locus.

The glycophorin locus (GYP) on the long arm of chromosome 4 encodes antigens of the MNSs blood group system and displays considerable allelic variation among human populations. The genomic structure and organization of a variant glycophorin allele specifying a novel Miltenberger (Mi)-related phenotype, MiX, were examined. This variant probably arose from a gene conversion event involving a direct repeat of the acceptor splice site. Southern blot analysis indicated that MiX gene derived its 5' and 3' portions from glycophorin B or delta gene but its internal part from glycophorin A or alpha gene. Genomic sequences encompassing the rearranged regions of the MiX gene were amplified by single copy polymerase chain reaction. Direct DNA sequencing showed that during the formation of MiX gene, a short stretch of alpha exon III with a donor splice site has replaced a silent sequence in the delta gene containing a cryptic acceptor splice site. The upstream delta-alpha breakpoint is flanked by the direct repeats of the acceptor splice site, whereas the down-stream alpha-delta breakpoint is located in the adjacent intron. This segmental transfer produced a new composite exon whose expression not only transactivated a portion of silent sequence but also created intraexon and interexon hybrid junctions that characterize the antigenic specificities of MiX glycophorin. The identification of MiX as yet another delta-alpha-delta hybrid different from MiIII and MiVI in gene conversion sites suggests that shuffling of expressed and unexpressed sequences through particular genomic DNA motifs has been an important mechanism for shaping the antigenic diversity of MNSs blood group system during evolution.

Alleles

The cold case of state transition 7 (stt7) mutants of Chlamydomonas reinhardtii, solved by whole-genome sequencing.

The process of State Transitions (ST) corresponds to an STT7 kinase-driven redistribution of the transmembrane LHCII antenna proteins between Photosystem II (PSII) and Photosystem I (PSI), which results from changes in their phosphorylation state. For the past two decades, two LHCII-kinase mutants, stt7-1 and stt7-9, have been instrumental in the study of STs in Chlamydomonas reinhardtii, the former being a null mutant for the kinase but quasi-sterile in crosses, while the latter, although fertile, has a leaky phenotype. Using long-read sequencing, this study further characterized the genetic lesions of the stt7 mutant strains through whole-genome reconstruction and de novo chromosome assembly. In addition, two new stt7 null mutants were generated, one derived by crosses from the original stt7-1 and one obtained by Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated protein 9 (Cas9) technology. This work provides a comprehensive genomic characterization of the original stt7-1 null mutant, revealing extensive chromosomal rearrangements and high levels of aneuploidy, associated with increased cell size and meiotic dysfunction. Reassessment of their physiology and genetic backgrounds highlights the need for caution in interpreting genetic information. We thus produced more reliable null mutants for the LHCII-kinase, amenable to genetic crosses for the study of STs in a variety of genetic backgrounds.

Chlamydomonas reinhardtii

Haplotype-resolved genome assembly and implementation of VitExpress, an open interactive transcriptomic platform for grapevine.

Haplotype-resolved genome assemblies were produced for Chasselas and Ugni Blanc, two heterozygous Vitis vinifera cultivars by combining high-fidelity long-read sequencing and high-throughput chromosome conformation capture (Hi-C). The telomere-to-telomere full coverage of the chromosomes allowed us to assemble separately the two haplo-genomes of both cultivars and revealed structural variations between the two haplotypes of a given cultivar. The deletions/insertions, inversions, translocations, and duplications provide insight into the evolutionary history and parental relationship among grape varieties. Integration of de novo single long-read sequencing of full-length transcript isoforms (Iso-Seq) yielded a highly improved genome annotation. Given its higher contiguity, and the robustness of the IsoSeq-based annotation, the Chasselas assembly meets the standard to become the annotated reference genome for V. vinifera. Building on these resources, we developed VitExpress, an open interactive transcriptomic platform, that provides a genome browser and integrated web tools for expression profiling, and a set of statistical tools (StatTools) for the identification of highly correlated genes. Implementation of the correlation finder tool for MybA1, a major regulator of the anthocyanin pathway, identified candidate genes associated with anthocyanin metabolism, whose expression patterns were experimentally validated as discriminating between black and white grapes. These resources and innovative tools for mining genome-related data are anticipated to foster advances in several areas of grapevine research.

Vitis

Complete plastid genome of Iris orchioides and comparative analysis with 19 Iris plastomes.

Iris is a cosmopolitan genus comprising approximately 280 species distributed throughout the Northern Hemisphere. Although Iris is the most diverse group in the Iridaceae, the number of taxa is debatable owing to various taxonomic issues. Plastid genomes have been widely used for phylogenetic research in plants; however, only limited number of plastid DNA markers are available for phylogenetic study of the Iris. To understand the genomic features of plastids within the genus, including its structural and genetic variation, we newly sequenced and analyzed the complete plastid genome of I. orchioides and compared it with those of 19 other Iris taxa. Potential plastid markers for phylogenetic research were identified by computing the sequence divergence and phylogenetic informativeness. We then tested the utility of the markers with the phylogenies inferred from the markers and whole-plastome data. The average size of the plastid genome was 152,926 bp, and the overall genomic content and organization were nearly identical among the 20 Iris taxa, except for minor variations in the inverted repeats. We identified 10 highly informative regions (matK, ndhF, rpoC2, ycf1, ycf2, rps15-ycf, rpoB-trnC, petA-psbJ, ndhG-ndhI and psbK-trnQ) and inferred a phylogeny from each region individually, as well as from their concatenated data. Remarkably, the phylogeny reconstructed from the concatenated data comprising three selected regions (rpoC2, ycf1 and ycf2) exhibited the highest congruence with the phylogeny derived from the entire plastome dataset. The result suggests that this subset of data could serve as a viable alternative to the complete plastome data, especially for molecular diagnoses among closely related Iris taxa, and at a lower cost.

Iris Plant

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

With hundreds of copies of rDNA, it is unknown whether they possess sequence variations that form different types of ribosomes. Here, we developed an algorithm for long-read variant calling, termed RGA, which revealed that variations in human rDNA loci are predominantly insertion-deletion (indel) variants. We developed full-length rRNA sequencing (RIBO-RT) and in situ sequencing (SWITCH-seq), which showed that translating ribosomes possess variation in rRNA. Over 1,000 variants are lowly expressed. However, tens of variants are abundant and form distinct rRNA subtypes with different structures near indels as revealed by long-read rRNA structure probing coupled to dimethyl sulfate sequencing. rRNA subtypes show differential expression in endoderm/ectoderm-derived tissues, and in cancer, low-abundance rRNA variants can become highly expressed. Together, this study identifies the diversity of ribosomes at the level of rRNA variants, their chromosomal location, and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Humans

Ancient DNA and Human Physiology.

Ancient DNA (aDNA) enables the reconstruction of chronologically sampled genomes from ancient humans, animals, plants, pathogens, and microorganisms, as well as environmental DNA, providing a record of biological changes through time. Improvements in short and degraded DNA extraction methods and low-cost sequencing now enable the generation of broad, cross-regional datasets that expand evolutionary analyses from past population demography to biological mechanisms. By tracking temporal shifts of allele frequencies, integrating functional genomics resources (e.g., gene expression, chromatin structure variation), modeling population demography to separate selection from genetic drift, and aligning genetic changes with archaeological, cultural, and climatic data, aDNA has the potential to link sequence variation to physiological function within their temporal and environmental contexts. In this review, we summarize illustrative case studies from aDNA research spanning complex traits, dietary adaptations, and responses to pathogens and other environmental changes, showing how human biology has evolved under multiple selective pressures through time. These dated signals help triage experimental work and expose mechanisms that are rare or absent in living cohorts. Although some challenges remain, such as geographic and temporal sampling disparities, limitations in data resolution and variant detection, and genotype-phenotype uncertainties, rapid methodological progress and stronger ethical frameworks are expanding what can be inferred, making aDNA a promising tool for refining physiological pathways, their timing, and their drivers.

Humans

Polyploidy-mediated variations in glutamate receptor proteins linked to Fusarium wilt resistance in upland cotton.

Cotton production in the US faces a serious threat from Fusarium oxysporum f. sp. vasinfectum race 4 (FOV4), a soil-borne fungus causing Fusarium wilt by infecting the roots and vascular system of susceptible cotton, leading to rapid wilting and death. Here, we investigate genetic mechanisms of resistance to FOV4 in the highly resistant upland cotton genotype "U1" using an early-generation segregating biparental population ("U1" × "CSX8308") with comprehensive genomic resources. Reference-grade genomic assemblies of the parents revealed minor structural variations between "U1" haplotypes, a high degree of collinearity at chromosome synteny and micro-synteny levels, and significant divergence from "CSX8308" with 8.9 million SNPs. QTL analysis identified significant markers on chromosomes D03 and A02 linked to reduced Fusarium wilt severity. Within these regions, two glutamate-receptor-like (GLR) genes showed structural variation and overlapped between translocated segments on A02 and D03, suggesting a rare but important reinforcing effect of parallel evolution between susceptible and resistant genotypes. Transcriptome profiles of "U1" under FOV4 infection reveal activation of calcium-binding proteins and transcription factors regulating plant hormones (ethylene, abscisic acid, jasmonic acid, and salicylic acid), along with enzymes involved in cell wall remodeling and phytoalexin production. Advancing cotton improvement depends on incorporating durable genetic disease resistance into high-yielding, high-quality cultivars.

Fusarium

Extensive longevity and DNA virus-driven adaptation in nearctic Myotis bats.

The genus Myotis is one of the largest clades of bats, and exhibits some of the most extreme variation in lifespans among mammals alongside unique adaptations to viral tolerance and immune defense. To study the evolution of longevity-associated traits and infectious disease, we generated cell lines and near-complete genome assemblies for 8 closely related species of Myotis. Using genome-wide screens of positive selection, analyses of structural variation, and functional experiments in primary cells, we identify new patterns of adaptation contributing to longevity, cancer resistance, and viral interactions in bats. We show that the recurrent evolution of longevity seen in Myotis leads to some of the highest predicted increases in cancer risk across mammals and demonstrate a unique DNA damage response in primary cells of the long-lived M. lucifugus. We also find evidence of abundant adaptation in response to DNA viruses - but not RNA viruses - in Myotis and other bats in sharp contrast with other mammals, potentially contributing to the role of bats as reservoirs of zoonoses. Together, our results demonstrate how genomics and primary cells derived from diverse taxa uncover the molecular bases of extreme adaptations in non-model organisms.

Aging

HLA-DP and HLA-DO genes in presumptive HLA-identical siblings: structural and functional identification of allelic variation.

We analyzed HLA class II genomic polymorphisms in three families in which bone marrow transplantation was performed between individuals presumed to be HLA identical, but in which unexplained mixed lymphocyte culture reactivity was observed. These families were characterized by classical HLA serology, MLC, and DP typing. In each family, a pair of "HLA-identical" siblings demonstrated a small proliferative response in bidirectional MLC. Southern blotting analysis performed with cDNA probes for DQ alpha, DP alpha, and DP beta identified DP genomic differences in each case. Hybridization of Bgl II-digested genomic DNA with a DP alpha cDNA probe revealed three prominent polymorphic fragments (7.7, 5.8, and 3.7 kb), which discriminated between presumptive identical siblings and indicated crossover events within HLA. Similarly, hybridization of SstI-digested genomic DNA with a DP beta cDNA probe, although resulting in a more complex pattern, identified DP genomic disparity between the presumed HLA identical siblings. Hybridization of SstI-digested DNA from two families with evidence of DP recombination was performed by using an oligonucleotide probe specific for the newly described HLA class II gene DO beta. Two major polymorphic fragments, at 6.2 and 3.3 kb, segregated in these families and localized the crossovers flanking the DO beta gene between the DQ and DP loci. The contribution of the antigenic differences marked by these HLA DP and DO DNA polymorphisms to allorecognition in MLR and in graft-vs-host disease are discussed.

Alleles

Plastid genome evolution and phylogenomics with broad taxon sampling: insights into intrafamilial classification of Hamamelidaceae.

Hamamelidaceae, within the order Saxifragales, comprises 27 genera and approximately 120 species. The family has a pantropical and temperate distribution across the Americas, Asia, Africa, and Australia. Previous molecular investigations, constrained by limited taxon sampling and inadequate genetic markers, supported a five-subfamily classification system. However, these studies predominantly focused on Asian taxa, resulting in poor resolution of the evolutionary relationships among American, African, and Australian genera. To address these sampling gaps, we employed near-complete generic sampling (26 of 27 genera) to investigate plastome architecture, structural variation, and phylogenetic relationships. We newly sequenced and assembled 15 plastid genomes representing geographically and taxonomically underrepresented genera and analyzed them alongside 59 publicly available plastomes retrieved from GenBank. Plastid genomes exhibited conserved quadripartite architecture with sizes ranging from 158, 076 bp to 160, 814 bp, minimal structural variation, consistent GC content (37.7-38.2%), and identical gene order. Inverted repeat (IR) regions had limited size variation (26, 211-26, 429 bp). Simple sequence repeat (SSR) distribution (2, 219 loci) showed no clear correlation with the genus-level phylogenetic relationships. We identified ten hypervariable regions, including coding sequences (accD, ycf1, clpP, ndhF, and rpl22) and intergenic spacers (rpl33-rps18, the trnG-UCC intron, trnH-GUG-psbA, accD-psaI, and petA-psbJ), as promising candidate regions for future applications in species delimitation and phylogenetic studies. Phylogenetic analyses revealed largely congruent topologies across datasets and methods, providing improved resolution and strong support for most subfamilial and tribal relationships compared with previous studies. This study highlights the utility of plastid genome data for resolving deep-level phylogenetic relationships within Hamamelidaceae. The genome architecture reflects the high conservation of plastid genomes, while the identified mutation hotspots represent potential resources for future taxonomic and phylogenetic studies. Our results support the existing subfamily classification while improving geographical coverage and generic representation, providing a robust framework for future taxonomic and evolutionary studies of this globally distributed and taxonomically complex family.

Hamamelidaceae

Comparisons Between Large-Scale Genomic Variants and SNPs in Driving Population Divergence and Local Adaptation.

Genomic variations, such as indels (2-49 bp) and structural variants (SVs, ≥50 bp), are larger-scale mutations than single nucleotide polymorphisms (SNPs) and can substantially impact evolutionary processes, including speciation, adaptation, and phenotypes. Despite their functional importance, integrative population genetic analyses that jointly consider genome-wide SNPs, indels, and SVs remain under-explored. The ground tit (Pseudopodoces humilis), an endemic species to the Qinghai-Tibet Plateau (QTP), exhibits divergence across distinct glacial refugia, accompanied by habitat and morphological divergence, making it an excellent example for investigating how different types of genomic variants contribute to population divergence and local adaptation. Here, by retrieving 81 whole-genome sequence data, over 13 million SNPs, 2 million indels, and 22,101 SVs were identified. Variants were unevenly distributed across the genome, characterized by distinct hotspot regions. Indels and SVs revealed four genetic clusters consistent with previous SNP-based results, thereby validating the reliability of our variant datasets. FST and genotype-environment association (GEA) analyses independently revealed numerous candidate indels and SVs; each showed minimal overlap with previously identified SNPs, and were enriched in similar functional pathways such as signal transduction, skeletal muscle development, water transport, DNA repair, reproduction, nervous system development, and immunity. Collectively, our results demonstrated that indels and SVs could capture additional signatures besides SNPs. Furthermore, similar but distinct gene functions among different types of genomic variants collectively and complementarily drive genomic divergence across environmental gradients in such a high-elevation endemic species, underscoring its evolutionary relevance in local adaptation.

indels

Dissecting genetic architecture and improving machine learning‑based genomic prediction of flowering time in Osmanthus fragrans by integrating structural variants.

Sweet osmanthus (Osmanthus fragrans), a traditional ornamental plant in China, exhibits substantial variation in autumn flowering time, which significantly affects landscape application and cultivation efficiency. Here, we performed a genome-wide association study on 127 resequenced accessions classified into early, intermediate, and late flowering types, using a set of 2,325,410 single-nucleotide polymorphisms (SNPs) and 246,824 structural variants (SVs). By integrating SNP/insertion and deletion (Indel) and SV data with weighted gene co-expression network analysis, machine learning, and genomic prediction, we dissected the genetic architecture of flowering time. We identified 24 associated SNP/Indels and six SVs, mapping to 30 candidate genes, including known flowering regulators FLK, LOS1, Y14, MIF2, and GID1B. These genes showed tissue-specific expression, with some responding to low temperature. The two hub genes, GUX1 and LYG027904, were located within modules of the co-expression network associated with low-temperature treatment. Haplotype analysis revealed a specific three-SNP haplotype associated with late flowering and linked to LOS1, and epistatic interactions among combined genotypes contributed to phenotypic variation. Notably, integrating SVs with SNP/Indels improved genomic prediction accuracy; the gradient boosting decision tree model outperformed other machine learning algorithms, achieving a mean accuracy of 0.859 and an AUC > 0.8 (where AUC is area under receiver operating characteristic curve) for all flowering types. These findings provide insights into the genetic mechanisms underlying flowering time variation in O. fragrans, offer candidate genes and haplotypes for molecular breeding, and highlight the value of integrating SVs with machine learning for genomic prediction in woody ornamentals.

Machine Learning