PubMed HealthSearch

Biomedical subjects

Genomics

Find indexed PubMed genomics citations. Search gene expression, sequencing and genetic variation in titles, abstracts and supplied subjects, then open the PubMed record.

At least 37 records · Page 2Linked to original sources

Mechanistic Perspectives From Genomics and Pangenomics of Medicinal and Aromatic Plants: Linking Genome Architecture to Phytochemical Diversity.

Medicinal and aromatic plants (MAPs) produce a remarkable diversity of specialized metabolites with significant pharmaceutical, nutraceutical, and industrial value. Although advances in long-read sequencing, chromosome-scale genome assembly, and pangenomics have greatly expanded genomic resources, the mechanistic links between genome architecture and phytochemical diversity remain incompletely understood. The present review synthesizes current evidence describing how structural genomic variation may contribute to phytochemical diversity, while acknowledging that many proposed genome-to-metabolite relationships require further experimental validation. Examples illustrate how genome architecture is associated with specialized-metabolite biosynthesis through multiple regulatory processes. However, the strength of supporting evidence varies considerably among MAP species. Moreover, relatively few genome-to-metabolite relationships have been confirmed through direct functional validation. We further discuss how pangenomics, multiomics integration, genome editing, synthetic biology, and artificial intelligence support the discovery, validation, and engineering of specialized metabolic pathways. Casual conclusions are evaluated according to the strength of available evidence, highlighting where causal relationships have been experimentally established and where conclusions remain primarily association-based. Overall, this review provides an integrated conceptual and evidence-based perspective summarizing proposed relationships between genome architecture and phytochemical diversity and outlines future priorities for functional genomics, precision breeding, metabolic engineering, and sustainable utilization of MAPs.

artificial intelligence

Low-pass whole-genome sequencing reveals genomic diversity and ecotype-specific adaptation in indigenous Tigrayan chickens.

Indigenous chickens play a critical role in food security and climate resilience in smallholder systems, yet their genomic diversity and adaptive potential remain insufficiently characterised. This study employed low-pass whole-genome sequencing (LP-WGS; 0.2-1.99×) to investigate genomic diversity, population structure, inbreeding and candidate environment-associated genomic variation in 33 chickens from highland, midland, and lowland agroecologies in the Tigray region of northern Ethiopia. After imputation and stringent filtering, 23.4 million high-confidence SNPs were retained, including ~ 17% novel variants, indicating substantial uncharacterised genetic diversity in these populations. SNP density (13.8 ± 8.6 SNPs/kb) was comparable to values reported from high-coverage Ethiopian chicken datasets, demonstrating the suitability of LP-WGS for population genomics in resource-limited settings. Marked differences in genomic diversity were observed among ecotypes: midland chickens showed the highest nucleotide diversity (π = 0.00267), followed by lowland (π = 0.00233), whereas highland chickens showed the lowest diversity (π = 0.00203) and elevated genomic inbreeding (FROH and FHOM ≈ 0.18). Population structure analyses revealed clear genetic separation among ecotypes. PCA (13.91% variation explained) distinguished lowland chickens along PC1 and separated highland from midland along PC2, while ADMIXTURE and FST patterns supported three major ancestral genomic backgrounds. Functional annotation of private missense variants uncovered distinct adaptive signatures reflecting the contrasting agroecological conditions. Highland chickens showed enrichment of candidate genes potentially involved in physiological processes relevant to high-altitude environments, including cold response, angiogenesis, cardiovascular regulation and metabolic homeostasis (eg., PARP1, ACOX2, ITGB3, EDNRB, SOX8, and SOX10). Midland chickens exhibited candidate signals of selection in genes with known roles in innate antiviral immunity, bacterial defence and inflammatory regulation (eg., BAK1, CLSTN1, CYSLTR1, CYSLTR2, CXCR7, GIPR, DSCAM, GDAP1, TLR3, TLR4, TLR7, IFIH1, ADORA1, EPHB1, and TMPRSS2). Lowland chickens displayed candidate variants associated with heat-stress response, DNA damage repair, oxidative balance and cardiovascular support under extreme temperatures (e.g., MLH1, BDKRB1, GPR19, FLT1, CCL18, TGM2, and RAMP3). Overall, the results indicate substantial genomic differentiation among ecotypes and suggest candidate environment-associated genetic divergence across Tigray's diverse agroecological zones. These populations may represent important reservoirs of adaptive genetic variation for climate-resilient poultry breeding, warranting further functional validation and conservation-oriented management.

Animals

Assembly and comparative analysis of the mitochondrial genome of Pleione yunnanensis: genome structure and evolutionary insights.

BACKGROUND: Pleione yunnanensis a terrestrial or semi-epiphytic herbaceous plant belonging to the Orchidaceae family, is valued for both its medicinal uses and ornamental appeal. Although its chloroplast genomes have been sequenced, its complete mt genome had not previously been resolved, limiting genetic and evolutionary studies of the species. RESULTS: In this work, we assembled and characterized the first complete mt genome of P. yunnanensis, revealing a structurally complex, multibranched system composed of 14 circular-mapping molecules totaling 468,176 bp with a GC content of 44.32%. The genome encodes 44 annotated genes, including 28 protein-coding genes (PCGs), 15 tRNAs, and one rRNA. The multibranched architecture provides new evidence supporting the dynamic and recombinational nature of plant mt genomes. Repeat analysis uncovered 29 simple sequence repeats (SSRs), 19 tandem repeats, and 118 dispersed repeats, indicating a comparatively lower repeat abundance than that found in closely related orchids with similar mt genome sizes. Codon-usage profiling of PCGs showed a marked bias toward A/T-ending codons. Prediction of RNA editing sites identified 4,708 putative edits across mitochondrial PCGs. Most mitochondrial genes displayed Ka/Ks ratios close to 1.0, suggesting relaxed selective constraints or lineage-specific evolutionary patterns rather than strong positive selection. Moreover, we detected 69 chloroplast-derived homologous fragments, including 15 intact genes, suggesting ongoing plastid-mitochondrial DNA transfer. Phylogenetic reconstruction and collinearity comparisons demonstrated that P. yunnanensis clustered closely with Dendrobium species, including D. amplum and D. hancockii, within the Orchidaceae clade. CONCLUSIONS: This study provides the first complete mt genome of P. yunnanensis, providing a foundational genomic resource for the genus Pleione. The results not only improve our understanding of mt genome structure and evolution in Orchidaceae, but also offer valuable molecular evidence for phylogenetic inference, germplasm identification, and conservation of this endangered medicinal species.

Orchidaceae

Highly Contiguous Is Not Chromosomally Accurate: Integrated Cytogenetic and Genomic Mapping in Two Turtle Genome.

High-quality genome assemblies are essential for robust research across biological and medical fields. Assembly errors can have far-reaching consequences for downstream analyses, including gene annotation and the inference of synteny. In contrast to the rapid growth of genomic data volume, there is a notable lag in the integration of chromosome-level assemblies with cytogenetic data. We conducted the first direct genome-to-genome comparison, integrating comparative chromosome painting, the alignment of chromosome-specific probes to available genome assemblies, and synteny-based comparison of independent chromosome-level assemblies of the loggerhead sea turtle (Caretta caretta, 2n = 56) and the red-eared slider (Trachemys scripta elegans, 2n = 50). Using two independent sets of flow-sorted chromosome-specific probes in cross-species hybridizations, together with the sequencing and mapping of chromosome-derived DNA libraries, we assigned assembled scaffolds to all physical chromosomes of both species. In C. caretta, chromosomal assignments and genome-wide synteny were fully consistent with the published assembly, except for the reduced sizes of two microchromosome scaffolds, which we attribute to under-representation of repetitive DNA. In contrast, in T. s. elegans, cytogenetic validation of the assemblies revealed a false rearrangement compared to a missed one. Our results show that even highly contiguous vertebrate genome assemblies can misrepresent chromosome structure. When cytogenetic analyses reveal such inaccuracies, updated reference genomes should be generated for widely studied species to enable accurate inference of karyotype evolution and downstream comparative genomic analyses.

FISH

Complete genome sequence and genomic characterization of the probiotic Limosilactobacillus reuteri PSC102.

BACKGROUND: Gut microbiota are potential sources of probiotics and play an essential role in maintaining intestinal health. Limosilactobacillus reuteri PSC102 (L. reuteri PSC102), which was isolated from the feces of healthy pigs, exhibited health-beneficial properties. AIM: We aimed to conduct a whole-genome sequencing analysis of L. reuteri PSC102 to determine its molecular characteristics as a probiotic strain. METHODS: Limosilactobacillus reuteri PSC102 cells were cultured in De Man-Rogosa-Sharpe medium, followed by DNA extraction for genomic analysis using the PacBio-Illumina sequencing platform. The EzBioCloud software was used to perform gene assembly, and the genes were interpreted by the National Center for Biotechnology Information (NCBI) and the Glimmer program. Core and pan-genomic analyses were performed to assess the extent of functional conservation in the genomic sequence. Moreover, the NCBI database and the Basic Local Alignment Search Tool software were used to identify antimicrobial resistance genes and virulence factors. RESULTS: Limosilactobacillus reuteri PSC102 consists of a single circular chromosome with 2,048,626 bp, a guanine- cytosine of 38.9%, 18 rRNA genes, and 69 tRNA genes. Among the 1,846 protein-coding sequences, genes associated with probiotic characteristics were identified, including genes involved in host-microbe interactions, stress tolerance, biogenesis, and defense mechanisms. Furthermore, the genome of L. reuteri PSC102 comprises 2,446 pan-genome and 1,222 core-genome orthologous gene clusters. A total of 74 unique genes were identified in L. reuteri PSC102 genome. These genes mostly encode proteins potentially involved in the transport and metabolism of amino acids and carbohydrates. Moreover, antibacterial resistance genes and virulence factors were absent in L. reuteri PSC102. CONCLUSION: The results of the molecular insight into L. reuteri PSC102 corroborates its use as a probiotic in humans and other animals.

Limosilactobacillus reuteri

Temporal Genomics Reveal a Century of Genomic Diversity Shifts Across a Biodiversity Hotspot Avian Assemblage.

Biodiversity has experienced tremendous shifts in community, species, and genetic diversity during the Anthropocene. Understanding temporal diversity shifts is especially critical in biodiversity hotspots, i.e., regions that are exceptionally biodiverse and threatened. Here, we use museomics and temporal genomics approaches to quantify temporal shifts in genomic diversity in an assemblage of eight generalist highland bird species from the Ethiopian Highlands (part of the Eastern Afromontane Biodiversity Hotspot). With genomic data from contemporary and historical samples, we demonstrate an assemblage-wide trend of increased genomic diversity through time, potentially due to improved habitat connectivity within highland regions. Genomic diversity shifts in these generalist species contrast with general trends of genomic diversity declines in specialist or imperiled species. In addition to genetic diversity shifts, we found an assemblage-wide trend of decreased realized mutational load, indicative of overall trends for potentially deleterious variation to be masked or selectively purged. Across this avian assemblage, we also show that shifts in population genomic structure are idiosyncratic, with species-specific trends. These results are in contrast with other charismatic and imperiled African taxa that have largely shown strong increases in population genetic structure over the recent past. This study highlights that not all taxa respond the same to environmental change, and generalists, in some cases, may even respond positively. Future comparative conservation genomics assessments on species groups or assemblages with varied natural history characteristics would help us better understand how diverse taxa respond to anthropogenic landscape changes.

Animals

Comparative analysis of chloroplast genomes in ten holly (Ilex) species: insights into phylogenetics and genome evolution.

In order to clarify the chloroplast genomes and structural features of ten Ilex species and provide insights into the phylogeny and genome evolution of the genus Ilex, we conducted a comparative analysis of chloroplast genomes using bioinformatics methods. The chloroplast genomes of ten Ilex species were obtained, and their structural features and variations were compared. The results indicated that all chloroplast genomes in the genus Ilex exhibit a double-stranded circular structure, with sizes ranging from 157,356 to 158,018 bp, showing minimal differences in size. The chloroplast genomes of the ten Ilex species have a relatively conservative gene count, with a total of 134 to 135 genes, including 88 or 89 protein-coding genes, and a conserved number of 8 rRNA genes. Each chloroplast genome contains 3 to 123 SSR (Simple Sequence Repeat) sites, predominantly composed of mononucleotide and trinucleotide repeats, with no detection of pentanucleotide or hexanucleotide repeats. The variation in dispersed repeat sequences among Ilex species is minimal, with a total repeat sequence number ranging from 1 to 14, concentrated in the length range of 30 to 42 base pairs. The expansion and contraction of chloroplast genome boundaries among Ilex species are relatively stable, with only minor variations observed in individual species. Variations in non-coding regions are more pronounced than those in coding regions, with the variability in the Large Single Copy region (LSC) being the highest, while the variability in the Inverted Repeat region A (IRa) is the lowest. The divergence time among Ilex species was estimated using the MCMC-tree module, revealing the evolutionary relationships among these species, their common ancestors, and their differentiation throughout the evolutionary process. The research findings provide a valuable reference for the systematic study and molecular marker development of Ilex plants.

Genome, Chloroplast

Complexity of schistosome vector bulinine snails in Kenya: Insights from nuclear genome size variation, complete mitochondrial genome sequence, and morphometric analysis.

Investigations of nuclear genome size, complete mitochondrial genome (mitogenome) sequence, and morphometrics were conducted on specimens of Bulinus snails (Gastropoda: Planorbidae) collected from 14 locations across the east coast, central Kenya, and western Kenya around the Lake Victoria region (November 2013 and January 2024). Flow cytometry measurements of DNA content (C-value) revealed unexpected variation in nuclear genome size, with diploid Bulinus africanus and B. forskalii species groups showing C-values ranging from 0.76 to 1.98 pg, while tetraploid B. truncatus had a C-value of 1.82 pg. Additionally, C-values for six B. globosus specimens from different localities ranged from 1.43 to 1.98 pg. These findings suggest that bulinine snails, particularly the B. africanus species group, have undergone genome expansion, whole genome duplication (polyploidization), or both, which have not been previously recognized. Next-generation sequencing was performed to determine and annotate 14 complete mitogenome sequences. Despite the well-conserved arrangement of protein-coding genes, two versions of mtDNA genome structure, distinguished by the tRNA-D (Asp) location, were found, designated as DCF (Asp-Cys-Phe) type (in the B. forskalii group and the B. truncatus/tropicus complex) and CF (Cys-Phe) type (in the B. africanus group). Phylogenetic analyses based on complete mtDNA sequences of bulinines from Kenya, along with cytochrome c oxidase subunit I (COX1) sequences from various localities across Africa, contributed to resolving species identities and provided further support for the presence of multiple or cryptic species in the taxon B. globosus. A landmark-based morphometric analysis was ineffective in distinguishing these species. This study reveals unexpected nuclear genome size variation, provides new mitogenome sequences, and highlights the limitations of morphological analysis. It offers valuable insights into the cytogenetics, polyploidy, genomics, taxonomy, and evolution of bulinines, which serve as intermediate hosts for schistosomes responsible for human urogenital schistosomiasis and intestinal schistosomiasis in domestic and wild mammals.

Animals

Revisiting the genome assembly of Lupinus species reveals differential diploidization after a shared whole-genome duplication.

Accurate genome assemblies are essential for comparative genomics, yet Hi-C-guided scaffolding can introduce structural errors that misrepresent chromosome architecture and bias evolutionary inferences. Here, we identified pervasive scaffolding errors-including artificial fusions, internal inversions, and incomplete contig mounting-in 2 previously published Lupinus genomes (L. cosentinii and L. digitatus) using a segmentation method based on long terminal repeat (LTR) retrotransposon density. We reassembled both genomes, producing chromosome-level references of 472.7 Mb (16 chromosomes) and 427.2 Mb (21 chromosomes), with BUSCO completeness >98.5%. Synteny validation and reapplication of LTR profiling confirmed that all prior errors were resolved. Using these corrected genomes together with 4 additional Lupinus species and 2 outgroup legumes, we investigated postpolyploid evolution. Synonymous substitution rate (Ks) analysis revealed a genus-specific whole-genome duplication (WGD) event (Ks = 0.17) shared by all 6 Lupinus species. The proportion of WGD-derived genes varied markedly, from 60% in L. digitatus to only 36% in L. mutabilis, indicating differential diploidization. While all species retained a core set of WGD duplicates enriched in cytoskeleton organization, ion transport, and defense responses, each exhibited lineage-specific functional trajectories: cell wall modification in L. cosentinii and L. digitatus, nitrogen metabolism in L. albus and L. angustifolius, flower development in L. luteus, and stress/lipid metabolism in L. mutabilis. Our corrected assemblies provide optimal references for Lupinus comparative genomics, and our findings demonstrate that a shared WGD event can lead to both conserved and highly divergent postpolyploid fates, likely underpinning adaptive diversification within the genus.

Lupinus

Comparative genomic analysis of Acer tsinglingense and A. davidii provides insights into nervonic acid biosynthesis, population evolution and genome vulnerability of endangered A. tsinglingense.

Global biodiversity is facing threats from climate change, habitat fragmentation, and anthropogenic activities-pressures that particularly endanger endemic and narrowly distributed species. In this study, the high-quality chromosome-level genomes of two ecologically divergent maples were assembled: the endangered and range-restricted Acer tsinglingense (791.40 Mb) and its widespread congener Acer davidii (1291.99 Mb). Phylogenomic analysis indicates that the two species diverged ~16.3 million years ago, with A. tsinglingense showing notable gene family expansions in secondary metabolite pathways. Notably, the 3-ketoacyl-CoA synthase gene family, which is involved in nervonic acid biosynthesis, underwent significant expansion and tandem duplication in A. tsinglingense, exhibiting high expression in buds. Population genomic analysis revealed that, compared with the widely distributed A. davidii, A. tsinglingense possesses lower genetic diversity, higher harmful mutation load, and signatures of a severe population bottleneck during the Late Pleistocene. Genome-environment association analysis further identified climate-adaptive genomic variations linked to five key environmental factors and projected potential genomic offsets under future climate scenarios. The southern lineage of A. tsinglingense exhibited greater climate sensitivity and genomic vulnerability under strong selective pressures, underscoring its importance as a conservation priority. Our research reveals that metabolic specializations in A. tsinglingense (such as the synthesis of nervonic acid) may confer competitive advantages in specific habitats. However, factors including its restricted distribution, historical population bottlenecks, and accumulated genetic load severely constrain its evolutionary potential to cope with rapid climate change. These findings emphasize the importance of elucidating the genomic basis and mechanisms of endangerment in metabolically specialized and threatened plant species to inform effective conservation strategies.

Genome, Plant

Genome-wide SNP data reveal geographic structure and landscape-associated genomic differentiation in a widespread lizard in arid Eastern Central Asia.

Arid landscapes provide important systems for examining how geographic structure and environmental heterogeneity shape genomic differentiation. In topographically complex desert regions, however, it remains challenging to determine whether population structure primarily reflects landscape resistance, geographic distance, or contemporary environmental variation. Here, we use genome-wide SNP data to investigate population structure, phylogenetic relationships, historical gene flow, demographic history, and landscape correlates of genomic differentiation in the variegated racerunner (Eremias vermiculata), a widespread lacertid lizard across arid Eastern Central Asia. Analyses of 164 individuals recovered six geographically structured nuclear clusters associated with major desert basins and mountain-bounded regions. Nuclear phylogenies resolved two broad regional clades corresponding to northeastern and southwestern parts of the species' range, while PCA and ADMIXTURE analyses recovered six finer-scale genetic clusters. Mitochondrial phylogenies, based on combined NCBI-derived Cyt b and COI sequences from the same individuals, recovered four deeper maternal lineages. These patterns indicate overall phylogeographic agreement between nuclear and mitochondrial datasets, with genome-wide SNPs providing finer-scale resolution of population structure. Demographic reconstructions further uncovered regionally heterogeneous Late Pleistocene histories among clusters, including signals of expansion, stability, and decline. Landscape genomic analyses revealed that genomic differentiation is primarily associated with landscape resistance, particularly elevation and land cover, as well as geographic distance, whereas contemporary environmental variables explained comparatively little variation after controlling for spatial structure. Together, our results suggest that genomic differentiation in E. vermiculata reflects the interplay of persistent landscape configuration, historical connectivity, and region-specific demographic histories across arid Eastern Central Asia. More broadly, this study highlights the value of integrating phylogeographic and landscape genomic approaches for understanding population differentiation and evolutionary history in topographically heterogeneous desert ecosystems.

Arid Eastern Central Asia

Chromosome-level genome assembly of the bitterling Rhodeus sinensis (Acheilognathidae) reveals genomic signatures associated with its mussel-dependent reproductive system.

Bitterlings (Acheilognathidae) exhibit a unique reproductive strategy characterized by symbiotic embryonic development inside the gill cavities of freshwater unionid mussels. Despite extensive ecological and physiological research on this system, genomic resources for bitterlings have remained limited, hindering comparative and evolutionary studies. Here, we present a high-quality, chromosome-level genome assembly for Rhodeus sinensis, a widely distributed bitterling species in the Korean Peninsula. By combining PacBio Continuous Long Read (CLR) sequencing, Illumina short reads, and Hi-C scaffolding, we generated a 0.77 Gb genome assembly with a scaffold N50 of 30.06 Mb. The final assembly comprises 24 chromosome-scale scaffolds, accounting for 98.3% of the assembled genome, with a BUSCO completeness score of 96.3% against the Actinopterygii_odb10. Comparative genomic analyses identified prominent expansions in gene families associated with alcohol metabolism, lipid catabolism, and oxidative stress responses. These genomic signatures of metabolic rewiring suggest a potential fuel flexibility, which may serve as a critical adaptive mechanism to mitigate the severe hypoxic stress encountered within the host mussel's gill environment. Ultimately, our chromosome-level genome assembly and findings provide a robust genomic foundation, contributing to a deeper understanding of the extreme physiological adaptations and unique life-history evolution within the Acheilognathidae.

Rhodeus sinensis

Comparative genomics of the monophasic variant of Salmonella Typhimurium: analysis of Colombian genomes and their relationship with international lineages.

The monophasic variant of Salmonella enterica serovar Typhimurium (STVM) represents a growing threat to global public health owing to its wide dissemination, capacity to adapt to multiple hosts, and antimicrobial resistance. In this study, 98 STVM isolates recovered in Colombia (57 from humans and 41 from pig farms and abattoirs) were genomically characterized between 2015 and 2022 and compared with 102 representative genomes of international lineages by whole-genome sequencing (WGS) and phylogenomic analysis. Phylogenomic analysis revealed the existence of two well-defined endemic lineages in Colombia (Clusters 1 and 2), arising from independent introduction events and subsequent local stabilization. Both lineages comprise isolates of human and swine origin without clear phylogenetic separation by host species, suggesting active zoonotic cocirculation and closely integrated interspecies transmission dynamics. Marked differences were observed in the accessory genome, including the differential presence of prophages (e.g., Gifsy-2, Fels-2, SW9), virulence plasmids, and resistance profiles. The Colombian lineages exhibited a high frequency of the pSTV plasmid (85%, n = 84/98) and a substantial burden of resistance determinants to quinolones (such as qnrB19, 74.5%; gyrA S83F mutation, 19.4%), phenicols (floR), tetracyclines (tetA, tetB), β-lactams (blaTEM-1B), and heavy metals. In contrast, the Colombian genomes clustered with the European ST34 lineage lacked pSTV but retained resistance and heavy metal operons. These findings demonstrate that international and endemic lineages coexist in Colombia with independent evolutionary trajectories, underscoring the need to strengthen genomic surveillance under the "One Health" approach to anticipate emerging threats and develop integrated control strategies.IMPORTANCEThe monophasic variant of Salmonella Typhimurium (STVM) has emerged as a predominant serovar in both humans and swine internationally. In Colombia, a fundamental question driving this study was whether local isolates belonged to international lineages or represented endemic strains. This study provides the first comprehensive genomic characterization demonstrating that two Colombian endemic lineages circulate simultaneously between humans and pigs without phylogenetic separation by host species, confirming active zoonotic transmission. The results demonstrate the coexistence of both lineages, each with distinctive repertoires of mobile genetic elements and specific antimicrobial resistance profiles. Understanding these transmission dynamics and evolutionary patterns is crucial for public health, as it demonstrates how zoonotic pathogens can establish locally adapted lineages with distinct resistance patterns. The genomic evidence of sustained interspecies circulation highlights the critical need for integrated surveillance strategies under the "One Health" framework. This will enable anticipating emerging threats, tracing transmission routes, and developing targeted interventions in food production systems.

One Health

Primulina pan-genome reveals differential gene retention following whole-genome duplications and provides insights into edaphic specialization.

Primulina, a genus of >200 species specialized to extreme soils, provides a model for edaphic adaptation. We assemble seven genomes and construct a pan-genome spanning nine species from karst, Danxia, and acidic soils. Comparative analyses reveal that karst-adapted species have smaller genomes. Two lineage-specific whole-genome duplications (WGDs) exhibit biased duplicate loss in large gene families but preferential retention of transcription factors, indicating combined adaptive and nonadaptive forces. Pan-genome analyses identify ion channel and transporter genes enriched in variant hotspots and under positive selection in karst lineages. Candidate genes for drought and salt stress tolerance include ABC transporters and ion channels. Notably, an ABC transporter shows positive selection in karst species and unique structural variation in non-karst species. Together, our findings show that genome downsizing, biased post-WGD retention, and evolution of ion-transport pathways shape adaptation to extreme soils. The Primulina pan-genome provides a resource for dissecting mechanisms underlying edaphic specialization.

Gene Duplication

Genomic Language Model for Predicting Enhancers and Their Allele-Specific Activity in the Human Genome.

Predicting and deciphering the regulatory logic of enhancers is a challenging problem, due to the intricate sequence features and lack of consistent genetic or epigenetic signatures that can accurately discriminate enhancers from other genomic regions. Recent machine-learning based methods have spotlighted the importance of extracting nucleotide composition of enhancers but failed to learn the sequence context and perform suboptimally. Motivated by advances in genomic language models, we developed DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. We trained two different models, using large collection of enhancers curated from the ENCODE registry of candidate cis-Regulatory Elements. The best fine-tuned model achieved 88.05% accuracy with Matthews correlation coefficient of 76% on independent set aside data. Further, we present the analysis of the predicted enhancers for all chromosomes of the human genome by comparing with the enhancer regions reported in publicly available databases. Finally, we applied DNABERT-Enhancer along with other DNABERT based regulatory genomic region prediction models to predict candidate SNPs with allele-specific enhancer and transcription factor binding activity. The genome-wide enhancer annotations and candidate loss-of-function genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies.

Journal Article

From family trials to genomic mate allocation: statistical and genomic strategies to accelerate sugarcane genetic improvement.

Sugarcane (Saccharum spp.) underpins global sugar and bioenergy supply and is increasingly valued as a renewable biomass feedstock. Sustained improvement in commercial traits and resilience is constrained by long breeding cycles, clonal propagation, multi-stage testing, and a highly polyploid, heterozygous, and frequently aneuploid genome with substantial non-additive genetic variation. Genomic selection has demonstrated value for predicting elite-clone performance, yet its operational use remains limited at earlier decision points, including family selection, parent evaluation, and cross design. This review examines the biological, statistical, and genomic factors that shape these decisions, with emphasis on the Australian breeding context based on progeny assessment trials (PATs), clonal assessment trials (CATs), and final assessment trials (FATs). We evaluate challenges arising from family plot means, the use of different full-sib samples as nominal family replicates, spatial heterogeneity, competition, genotype-by-environment interaction, and the partitioning of additive and non-additive effects. We also assess the integration of pedigree and genomic relationship, genotype representation, allele-dosage estimation, aneuploidy, genomic prediction models, and training-population design. We then consider genomic prediction of cross performance and constrained mate allocation as approaches for improving expected family performance, accounting for cross-specific non-additive effects and managing relatedness. We propose a decision-centred framework that links family and clonal data across breeding stages, tracks the propagation of information and uncertainty, and supports parent recycling and cross allocation. We conclude with a practical research agenda for stage-integrated mixed-model and single-step analyses that connect early family evaluation with genomic prediction and cross-level decision support in sugarcane breeding.

Saccharum

Genome editing research initiatives and regulatory landscape of genome edited crops in India.

Food and nutritional security are the top priorities in Indian agriculture. Exponential population growth coupled with climate change effects has become a serious challenge for sustainable agriculture. Genome editing has revolutionized the agricultural sector because of its ability to create precise, stable and predictable modifications in the genome and therefore, offers great opportunities for crop improvement in India. However, for harvesting the real benefits of this technology in agriculture sector, there is a strong need of creating awareness among the end users and development of suitable policies for regularization of genome edited products. Many regulatory agencies around the world have been modernizing their regulatory approaches to be more risk proportionate and to reflect a more science-based approach. In this article, recent research initiatives and developments undertaken by different Indian institutes/organizations for the genetic improvement of agricultural and horticultural crops via genome editing technologies are summarized. Furthermore, to benefit from this potential technology in our country, regulatory policies must be clear, science-based and proportionate. Therefore, in the present review, the regulatory policies related to the genome editing of crop products in India are discussed in detail. This review will sensitize researchers and stakeholders to the application of genome editing techniques in crop improvement and various biosafety committees involved in the development and regulation of genome edited crops.

Crops, Agricultural

Genomic prediction and genome-wide association study for liver abscesses in crossbred beef cattle.

Liver abscesses are a concern in feedlot cattle, and little is known about the role of genetics in their development. This study aimed to estimate genetic parameters and to identify single-nucleotide polymorphisms (SNPs) associated with liver abscesses. Crossbred cattle representing 18 breeds in the U.S. Meat Animal Research Center Germplasm Evaluation Program were phenotyped for liver abscesses at slaughter (n&#x2005;=&#x2005;9,044). Seventeen percent of cattle had liver abscesses. These cattle had genotypes that were imputed to sequence variant genotypes. After filtering and quality control, 340,723 SNPs were used in the analysis. Liver abscess prevalence was modeled with a single-step genomic best linear unbiased prediction (ssGBLUP) threshold model using a Bayesian framework. The model included contemporary group (sex, treatment group, and slaughter date), additive genomic, and residual effects. Genomic heritability was 0.039 (95% highest posterior density&#x2005;=&#x2005;0.005, 0.081), which was very small. To assess prediction quality, a 5-fold random cross-validation structure was used. Method Linear Regression was used to assess accuracy, bias, and dispersion by comparing estimated breeding values (EBV) from full and reduced analyses. Cross-validation metrics showed EBV based on genotypes had 0.05 reliability (SD&#x2005;<&#x2005;0.01) with no bias relative to EBV based on genotypes and phenotypes. For the genome-wide association study, SNP effects were back calculated from the EBV solutions from ssGBLUP. No SNPs were associated with liver abscesses at a Benjamini-Hochberg adjusted 0.05 significance level. Although a large dataset was used, this result was because of the low genomic heritability and imprecise EBV used to calculate SNP effects. Based on these results, environmental factors contribute to most of the variation in liver abscesses. Genetic selection to reduce liver abscesses would be slow because of the low genomic heritability, measurement late in life, and inability to measure breeding animals. A faster approach would be finding additional environmental interventions that maintain animal performance.

Animals