PubMed HealthSearch

SEARCH · PubMed Health

Results for “Intron annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Long-read transcriptomics corrects Trichomonas vaginalis intron annotations and refines transcript-end features.

BACKGROUND: Trichomonas vaginalis causes the most prevalent non-viral sexually transmitted infection worldwide. Despite its large genome (181.5 Mb; 36,310 predicted protein-coding genes in NYU_TvagG3_2), intron annotations remain limited and inconsistently validated. A recent short-read RNA-seq study reported 63 putative active introns, but short reads can misassign splice boundaries and cannot resolve complete transcript structures. METHODS: We integrated Oxford Nanopore direct RNA sequencing (DRS), ONT cDNA long-read sequencing, and Illumina RNA-seq to refine intron annotations, transcript-end features, and UTR boundaries in T. vaginalis. Candidate introns were validated by targeted PCR and Sanger sequencing, and representative splicing events were further assessed using public SRA datasets. RESULTS: Starting from 31 historically annotated introns, motif-guided long-read screening and orthogonal validation identified 17 additional validated introns, increasing the curated set to 48 confirmed introns. Among these 17 events, three were previously unrecognized in the current NYU_TvagG3_2 reference annotation. We also corrected five reported loci, including two false-positive introns, two splice-coordinate misannotations, and one gene-sequence error. DRS further supported transcript termination site mapping, UAAA polyadenylation-signal profiling relative to poly(A) addition sites, and single-molecule poly(A)-tail estimation. StringTie mixed-mode assemblies provided updated UTR boundaries for intron-bearing transcripts and transcripts without curated introns. CONCLUSIONS: This study provides a rigorously validated, long-read-refined resource of intron annotations, UTR boundaries, and UAAA-guided transcript-end features for T. vaginalis, together with a reproducible workflow for non-model protists. These refinements improve the current reference annotation and support future studies of functional genomics, parasite biology, pathogenesis, and diagnostic development.

Trichomonas vaginalis

Whole-genome sequencing in 333,100 individuals reveals rare non-coding single variant and aggregate associations with height.

The role of rare non-coding variation in complex human phenotypes is still largely unknown. To elucidate the impact of rare variants in regulatory elements, we performed a whole-genome sequencing association analysis for height using 333,100 individuals from three datasets: UK Biobank (N&#x2009;=&#x2009;200,003), TOPMed (N&#x2009;=&#x2009;87,652) and All of Us (N&#x2009;=&#x2009;45,445). We performed rare (&#x2009;<&#x2009;0.1% minor-allele-frequency) single-variant and aggregate testing of non-coding variants in regulatory regions based on proximal-regulatory, intergenic-regulatory and deep-intronic annotation. We observed 29 independent variants associated with height at P&#x2009;<&#x2009;after conditioning on previously reported variants, with effect sizes ranging from -7cm to +4.7&#x2009;cm. We also identified and replicated non-coding aggregate-based associations proximal to HMGA1 containing variants associated with a 5&#x2009;cm taller height and of highly-conserved variants in MIR497HG on chromosome 17. We have developed an approach for identifying non-coding rare variants in regulatory regions with large effects from whole-genome sequencing data associated with complex traits.

Humans

Beyond exons: Linking noncoding heritability and polygenicity across complex human traits and disorders.

The genetic architecture of complex traits spans a continuum of polygenicity, yet it remains unclear how differences in polygenicity relate to the functional localization of SNP heritability across the genome. We use a MiXeR-based framework to partition heritability across 74 functional annotations covering exonic, intronic, and intergenic regions for 34 complex traits and introduce a likelihood-based annotation contribution score that quantifies annotation-specific impact on heritability. Exons account for a minority of heritability, and their contribution decreases with increasing polygenicity, from an average of 22% in less-polygenic somatic diseases and biomarkers to 13% in highly polygenic psychiatric and cognitive phenotypes. Intergenic fractions show the opposite trend, whereas intronic fractions remain relatively stable. Analysis of the broader set of functional annotations also reveals systematic differences along the polygenicity axis: highly polygenic traits show stronger contributions from comparative genomics and variant-effect scores, whereas less-polygenic traits show stronger contributions from promoter, transcription, and chromatin annotations. Together, these results indicate that the functional partitioning of heritability systematically varies with polygenicity, shifting from gene-proximal regulatory architectures to architectures shaped by numerous dispersed regulatory effects.

MiXeR

Beyond Exons: Linking Noncoding Heritability and Polygenicity across Complex Human Traits and Disorders.

The genetic architecture of complex traits spans a continuum of polygenicity, yet it remains unclear how differences in polygenicity relate to the functional localization of SNP heritability across the genome. We use a MiXeR-based framework to partition heritability across exonic, intronic, and intergenic regions for 34 traits and introduce a likelihood-based annotation contribution score that quantifies annotation-specific impact on heritability. Exons explain a minority of heritability, and their contribution decreases with increasing polygenicity, from an average of 22% in less polygenic somatic diseases and biomarkers to 13% in highly polygenic psychiatric and cognitive phenotypes. Intergenic fractions show the opposite trend, whereas intronic fractions remain relatively stable. Analysis of a broader set of functional annotations reveals systematic differences along the polygenicity axis: highly polygenic traits show stronger contributions from comparative genomics and variant-effect scores, whereas less polygenic traits show stronger contributions in promoter, transcription, and chromatin annotations. Together, these results indicate that the functional partitioning of heritability systematically varies with polygenicity, pointing to a shift from gene-proximal regulatory architectures to architectures shaped by numerous dispersed regulatory effects as a key determinant of differences in polygenicity across traits.

Journal Article

Circular RNA profiling reveals an abundant circLMO7 that regulates myoblasts differentiation and survival by sponging miR-378a-3p.

Circular RNAs (circRNAs) have been identified from various tissues and species, but their regulatory functions during developmental processes are not well understood. We examined circRNA expression profiles of two developmental stages of bovine skeletal muscle (embryonic and adult musculus longissimus) to provide first insights into their potential involvement in bovine myogenesis. We identified 12&#x2009;981 circRNAs and annotated them to the Bos taurus reference genome, including 530 circular intronic RNAs (ciRNAs). One parental gene could generate multiple circRNA isoforms, with only one or two isoforms being expressed at higher expression levels. Also, several host genes produced different isoforms when comparing development stages. Most circRNA candidates contained two to seven exons, and genomic distances to back-splicing sites were usually less than 50&#x2009;kb. The length of upstream or downstream flanking introns was usually less than 105&#x2009;nt (mean&#x2248;11&#x2009;000&#x2009;nt). Several circRNAs differed in abundance between developmental stages, and real-time quantitative PCR (qPCR) analysis largely confirmed differential expression of the 17 circRNAs included in this analysis. The second part of our study characterized the role of circLMO7-one of the most down-regulated circRNAs when comparing adult to embryonic muscle tissue-in bovine muscle development. Overexpression of circLMO7 inhibited the differentiation of primary bovine myoblasts, and it appears to function as a competing endogenous RNA for miR-378a-3p, whose involvement in bovine muscle development has been characterized beforehand. Congruent with our interpretation, circLMO7 increased the number of myoblasts in the S-phase of the cell cycle and decreased the proportion of cells in the G0/G1 phase. Moreover, it promoted the proliferation of myoblasts and protected them from apoptosis. Our study provides novel insights into the regulatory mechanisms underlying skeletal muscle development and identifies a number of circRNAs whose regulatory potential will need to be explored in the future.

Animals

Efficient evidence-based genome annotation with EviAnn.

For many years, machine learning-based ab initio gene finding approaches have been central components of eukaryotic genome annotation pipelines, and they remain so today. The reliance on these approaches was originally sustained by the high cost and low availability of gene expression data, a primary source of evidence for gene annotation along with protein homology. However, innovations in modern sequencing technologies have revolutionized the acquisition of gene expression data, allowing scientists to rely more heavily on this class of evidence. In addition, proteins found in a multitude of well-annotated genomes represent another invaluable resource for gene annotation. Existing annotation packages often underutilize these data sources, which prompted us to develop EviAnn (Evidence-based Annotator), a novel evidence-based eukaryotic gene annotation system. EviAnn takes a strongly data-driven approach, building the exon-intron structure of genes from transcript alignments or protein-sequence homology rather than from purely ab initio gene finding techniques. We show that when provided with the same input data, EviAnn consistently outperforms current state-of-the-art packages including BRAKER3, MAKER2, and FINDER, while utilizing considerably less computer time. Annotation of a mammalian genome can be completed in less than an hour on a single multi-core server. EviAnn is freely available under an open-source license from https://github.com/alekseyzimin/EviAnn_release and from Bioconda as "eviann".

Journal Article

Pangenome-wide identification and expression analysis of the chalcone synthase (CHS) gene family in five yellowhorn spp.

Chalcone synthase (CHS) is a pivotal enzyme in flavonoid biosynthesis involved in plant development, defense, and secondary metabolism. Xanthoceras sorbifolium (yellowhorn) is a medicinal and ornamental species with high resistance to environmental stresses, but its CHS gene family remains uncharacterized. We performed a pangenome-wide identification of CHS genes across five yellowhorn genomes (Xzs4, Xwf8, Xjg, Xg11, and Xzg2). Across the five yellowhorn genomes, 27 CHS genes were identified and classified into four core pangenes, present in all five genomes, and two dispensable genes, present only in a subset of genomes. Phylogenetic analysis grouped these genes into three major clades, and chromosomal mapping and duplication analyses identified four tandemly duplicated gene pairs under purifying selection. The analyses of conserved structural features, including protein motifs and exon-intron organization, together with promoter cis-regulatory elements and gene ontology annotation, further indicated the potential involvement of CHS genes in flavonoid biosynthesis and stress-responsive mechanisms. Gene expression profiling identified significant upregulation of Xg11_CHS1 and Xg11_CHS3 under cold and drought stress, with tissue-specific expression patterns. These findings provide valuable insights into the evolution, functional diversification, and stress-responsive roles of the CHS gene family, identifying candidate genes for future studies targeting stress tolerance and flavonoid biosynthesis in yellowhorn.

Acyltransferases

Fishing for a reelGene: evaluating gene models with evolution and machine learning.

Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.

Machine Learning

The Role of Small Segmental Duplications in Generating Identical Isoforms Through Alternative Splicing Sites.

Alternative splicing plays a crucial role in expanding proteomic diversity but can also generate identical isoforms under certain conditions. While mutually exclusive splicing of tandem exons has occasionally been reported to produce identical isoforms, the extent to which other splicing events contribute to this phenomenon remains unclear. In this study, we demonstrate that alternative 5' and 3' splice site selection can also lead to the formation of identical isoforms, providing an additional type of splicing event for functional redundancy in transcriptomes. To address this, we analyzed reference genome annotations from 15 plant species, including Arabidopsis thaliana and wheat (Triticum aestivum), obtained from the RefSeq database. Identical isoforms were computationally defined as transcripts with distinct exon-intron structures but identical coding sequences. Our analysis reveals that the majority of alternative 5' and 3' fragments originate from small segmental duplications, suggesting that sequence repetition within gene regions facilitates the emergence of such splicing patterns. We also observed differences in the annotated 5' UTRs of some identical isoforms. However, since the alternative splicing sites themselves were not located within UTRs, these differences may reflect annotation uncertainty rather than genuine AS-derived variation. Given that UTR predictions in reference databases are not always precise, such observations should be interpreted cautiously. Expression analysis using an isoform-specific k-mer approach confirmed that identical isoforms can be differentially regulated. These findings suggest that, beyond expanding protein diversity, alternative splicing can also generate redundant isoforms that are differentially expressed at the RNA level, indicating potential regulatory roles. By elucidating the structural and regulatory factors contributing to the formation and retention of identical isoforms, our study provides new insights into the evolutionary and functional significance of alternative splicing in plants.

Alternative Splicing

Dissecting the genetic basis underlying drought tolerance at different development stages in soybean.

INTRODUCTION: Soybean is an indispensable crop supplying protein and oil for humans and animals, and playing an essential role in global food security. Drought represses soybean seed germination, reducing biomass accumulation and even inhibiting yield. METHODS: In order to dissect the genetic components underlying soybean drought tolerance during different development stage, a natural population containing 140 accessions was employed to evaluate seven drought tolerance-related traits under water-welled and drought stress conditions. Subsequently, genome-wide association study (GWAS) was conducted based on 150K single nucleotide polymorphism (SNP) markers of "Zhongdouxin-1". And the drought tolerance coefficient of seven different traits were analyzed with seven GWAS models. RESULTS: A total of 1807 significant SNPs were detected across 20 chromosome, including 569 SNPs for germination stage, and 1242 SNPs for seedling stage. Of 569 SNPs identified in germination stage, 354 SNPs on chromosomes 2, 7, 13, 14, and 17 accounting for 62.21%. Among 1242 SNPs found in seedling stage, 869 SNPs on chromosomes 11, 14, 15, 17 and 18 accounting for 69.97%. Moreover, among 1807 significant SNPs, 163 SNPs exhibited pleiotropic effects, of which 23 were located in exon, 21 in intron, 12 in 5'UTR or 3'UTR and 11 in upstream or downstream. Furthermore, 249 stable SNPs were detected by more than four GWAS models. According to these stable SNPs, RNA expression levels and gene annotations, four causal genes (Glyma.02G080200, Glyma.11G056200, Glyma.12G188900, and Glyma.18G110200) conferring soybean drought tolerance were detected, which participated in ethylene stimulus response, water deprivation response, and proteolysis. DISCUSSION: Collectively, 249 stable SNPs, 163 pleiotropic SNPs and four candidate genes identified in present study provided promising molecular resources and reliable foundation for drought resistance improvement and marker-assisted selective breeding in soybean.

GWAS

Bayesian reconstruction and differential testing of excised introns.

MOTIVATION: Characterizing the differential excision of introns is critical for understanding the functional complexity of a cell or tissue, from normal developmental processes to disease pathogenesis. Most transcript reconstruction methods infer full-length transcripts from high-throughput sequencing data. However, this is a challenging task due to incomplete annotations and the heterogeneous expression of transcripts across cell-types, tissues, and experimental conditions. Several recent methods circumvent these difficulties by considering local splicing events, but these methods lose transcript-level splicing information and may conflate similar, but distinct transcripts. RESULTS: In this work, we formalize a new transcript reconstruction problem that interpolates between the full-length and local splicing perspectives by considering sequences of exon-exon junctions (SEEJs) that co-occur in transcripts. We then present a hierarchical Bayesian admixture model and posterior inference algorithms for computing SEEJs (BSEEJ), and a generalized linear model for characterizing differential SEEJ usage based on model parameter estimates. We show that BSEEJ achieves high F1 score for reconstruction tasks and improved accuracy and sensitivity in differential splicing when compared with six transcript and local splicing methods on simulated data. Lastly, we evaluate BSEEJ on experimental data based on transcript reconstruction, novelty of transcripts produced, model sensitivity to hyperparameters, and a functional analysis of differentially expressed SEEJs. AVAILABILITY AND IMPLEMENTATION: BSEEJ is freely available at https://github.com/bayesomicslab/BSEEJ.

Bayes Theorem

Tackling non-canonical splicing in arrhythmogenic cardiomyopathy to reduce the uncertain significance variants burden.

BACKGROUND: Splice-altering variants (SAVs), particularly those outside canonical splice sites, are an underappreciated contributor to inherited cardiovascular diseases. In arrhythmogenic cardiomyopathy (ACM), these variants frequently remain classified as of uncertain significance (VUS) due to limited predictive power and lack of transcript-level evidence, constraining genetic yield and clinical management. Our study aimed to determine the functional impact of SAVs in ACM genes and refine their classification using ACMG/AMP and ClinGen SVI criteria. METHODS: SAVs identified in 200 ACM probands underwent SpliceAI prediction, GTEx cardiac exon-usage annotation, and functional assessment using pSPL3-based minigene assays. Aberrant transcripts were quantified using Percent Splicing Alteration (PSA). Segregation data and ACMG/AMP criteria refined by ClinGen SVI were applied to integrate functional and clinical evidence for classification. RESULTS: Aberrant splicing was confirmed in 9/20 variants (45%), including synonymous, missense, and non-canonical intronic changes. SpliceAI scores correlated strongly with PSA values (R&#xb2;=0.86). Case-control burden testing revealed significant enrichment of splice-altering variants in DSP, DSG2, DSC2 and FLNC. Integrating predictive algorithms with experimental validation and segregation analysis markedly enhances reclassification of 16/20 variants (80%). CONCLUSION: Splicing defects beyond canonical sites significantly shape ACM genetic landscape. Integrating predictive models with experimental validation clarifies uncertain variants bridging the gap between genomic uncertainty and clinical decision-making.

Humans

Mitochondrial genomic characteristics and phylogenetic analysis of Cunninghamella elegans (Mucorales: Cunninghamellaceae).

Cunninghamella, a filamentous fungal genus with important biomedical and biochemical value, lacks any fully annotated mitochondrial genome to date. Herein, we presented the first complete mitogenome of Cunninghamella elegans, a circular 41,552 bp molecule (GC 27.86%) encoding 14 conserved protein-coding genes, 2 rRNA genes, 24 tRNA genes, and 6 non-conserved ORFs. Structural comparison with related species (Absidia glauca and Gongronella sp. w5) revealed dynamic evolution in intron and repeat elements. Phylogenetics places C. elegans within Cunninghamellaceae, with Gongronella as its closest relative. This reference mitogenome will underpin future evolutionary and taxonomic investigations of this industrially and medically significant lineage.

Cunninghamella elegans

The complete and annotated mitochondrial genome of Hemileia vastatrix Race I, causal agent of coffee leaf rust.

Hemileia vastatrix is the fungal pathogen responsible for coffee leaf rust (CLR), the most economically important disease of Coffea arabica worldwide. Recently, the nuclear genome of this fungus was completely deciphered. However, the mitochondrial genome of H. vastatrix has remained undercharacterized. Here, we present the complete, circularized mitochondrial genome of H. vastatrix Race I (isolate HvRI), assembled using a hybrid approach combining PacBio HiFi long reads and BGIseq short reads. The genome is 173,525&#xa0;bp in length with a GC content of 33.1% and encodes 41 functional genes, including 15 protein-coding genes, 2 rRNAs, and 24 tRNAs. The assembly reveals significant structural complexity, driven by intron expansion in the cox1 and cob genes. Notably, the atp8 gene contains a group II intron, rare for this locus, whose internal open reading frame displays evidence of pseudogenization via internal stop codons.. We also characterized a putative replication initiation zone (~1.2&#xa0;kb) defined by a poly-G homopolymer and conserved regulatory motifs. The mitogenome of the HvRI isolate does not contain cob mutations that lead to amino acid substitutions G143A and F129L associated with the quinone outside inhibitor (QoI) fungicide resistance. This high-quality mitogenome is an important resource for comparative mitogenomics, population diversity studies, and the molecular surveillance of QoI fungicide resistance.

Genome, Mitochondrial

Mitochondrial genome characteristics and phylogenetic analysis of Ramaria longispora.

This study, for the first time, assembled and annotated the complete mitochondrial genome of R.&#xa0;longispora using high-throughput sequencing technology. The genome is a circular molecule with a total length of 157,712&#x2009;bp and a GC content of 31.55%. It encodes 71 genes, including 15 core protein-coding genes (PCGs), 25 transfer RNA (tRNA) genes, 2 ribosomal RNA (rRNA) genes, 5 free-stranding open reading frames (ORFs), and 24 intronic ORFs. Among these, most free-stranding ORFs have unknown functions but include a DNA polymerase gene, while the intronic ORFs primarily encode LAGLIDADG and GIY-YIG endonucleases. The mitochondrial genome contains 39 introns. Phylogenetic analyses based on 15 core PCGs using Bayesian inference (BI) and maximum likelihood (ML) methods revealed that this R. longispora is most closely related to Ramaria flavescens and Ramaria ichnusensis. This study provides foundational data for mitochondrial genome research in the Ramaria genus and offers important references for taxonomic and evolutionary studies of this group.

Mitochondrial genome

Dual Aberrant Splicing Caused by an Apparently Missense CHD7 Variant, c.5273A>G (p.Asp1758Gly), in CHARGE Syndrome.

CHARGE syndrome is a rare congenital disorder primarily attributed to heterozygous pathogenic variants of the CHD7 gene. Most pathogenic CHD7 variants are loss-of-function (LoF) variants, whereas the interpretation of missense variants remains challenging in the absence of functional evidence for their pathogenicity. We report a female infant presenting with clinical features characteristic of CHARGE syndrome. Targeted sequencing identified a heterozygous CHD7 variant (NM_017780.4:c.5273A>G), initially annotated as a missense substitution p.Asp1758Gly. This variant has been previously reported and registered with conflicting pathogenicity classifications; however, its transcript-level consequences remain unclear. Long-PCR-based RNA sequencing of total RNA from peripheral blood mononuclear cells revealed two aberrant splicing patterns associated with the variant: a predominant transcript carrying a 28-bp deletion due to cryptic donor splice-site activation, and a minor transcript with partial intron 24 retention. Both transcripts were predicted to result in premature termination codons. These findings demonstrate that c.5273A>G functions as a LoF variant through dual aberrant splicing rather than a simple missense substitution. This case underscores the importance of RNA-level splicing analysis for the accurate interpretation and classification of CHD7 missense variants.

CHD7

Improving spliced alignment by modeling splice sites with deep learning.

MOTIVATION: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands sophisticated modeling of splice sites, but current aligners use simple models, which may affect their accuracy given dissimilar sequences. RESULTS: We implemented minisplice to learn splice signals with a one-dimensional convolutional neural network (1D-CNN) and trained a model with 7,026 parameters for vertebrate and insect genomes. It captures conserved splice signals across phyla and reveals GC-rich introns specific to mammals and birds. We used this model to estimate the empirical splicing probability for every GT and AG in genomes, and modified minimap2 and miniprot to leverage pre-computed splicing probability during alignment. Evaluation on human long-read RNA-seq data and cross-species protein datasets showed our method greatly improves the junction accuracy especially for noisy long RNA-seq reads and proteins of distant homology. AVAILABILITY AND IMPLEMENTATION: https://github.com/lh3/minisplice.

Journal Article

Development and validation of a high-density 'Amahysnp' genotyping array in grain amaranth (Amaranthus hypochondriacus).

BACKGROUND: Grain amaranth has recently gained global attention as a promising crop alternative to traditional cereals due to its nutritional value and adaptability to various growing conditions. Although gene banks conserve extensive collections of amaranth germplasm, the genomic and phenotypic characterization of these resources is limited, which hinders their full utilization in breeding programs. A major challenge is the lack of high-throughput genotyping assays essential for comprehensive genomic characterization and trait mapping. High-density SNP arrays have become standard tools for genome-wide analysis across multiple loci, enabling molecular breeding across a range of crop species. RESULTS: In this study, we developed a 64&#xa0;K high-throughput SNP genotyping array named "AmahySNP", using Affymetrix&#xae; Axiom&#xae; technology. The array contains 64,069 high-density SNPs distributed across both genic (55.17%) and non-genic (44.83%) regions of the Amaranthus hypochondriacus genome. The genic region includes 8,879 genes, which consist of 4,830 single-copy genes and 4,049 multi-copy genes distributed across 16 scaffolds. These genes cover various functional regions, including exons (10.5%), introns (40.1%), 5'UTRs (1.6%), and 3'UTRs (2.9%), respectively. The AmahySNP array was effectively utilized for population structure analysis, genetic diversity studies, core development, and genome wide association studies (GWAS) in amaranth germplasm. A representative core set of 112 accessions was identified, which includes two released varieties (Annapurna and Suvarna) and 100 diverse accessions from 12 different regions, representing 12% of the total 917 accessions evaluated. Phylogenetic analysis revealed three major genetic clusters, independent of their geographical origins. GWAS conducted using 22,763 polymorphic SNPs from 540 genotypes identified 13 novel loci associated days to flowering (DTF) trait, seven of which were located within annotated genes. CONCLUSIONS: The AmahySNP 64&#xa0;K SNP chip a valuable genomic tool for amaranth research and breeding with a strong potential to accelerate its genetic improvement. It enables high-throughput genotyping for a wide range of applications, including GWAS and other genomic studies, and will significantly advance the exploration of natural genetic variations. Ultimately, this resource will empower amaranth breeders to develop improved amaranth cultivars with enhanced crop yield, resilience, and nutritional quality, contributing to global food security and sustainable agriculture.

Amaranthus