PubMed HealthSearch

SEARCH · PubMed Health

Results for “Simple repeated sequences”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Assembly and characterization of the first complete mitochondrial genome of Epimedium sagittatum (Sieb. et Zucc.) Maxim (Berberidaceae):an invaluable traditional Chinese medicine.

BACKGROUND: Epimedium sagittatum (Sieb. et Zucc.) Maxim is an invaluable traditional Chinese medicine plant known for its properties of tonifying kidney yang, strengthening bones and muscles, and dispelling rheumatism. The chloroplast (cp) genome of E. sagittatum have been sequenced, offering critical insights for breeding and phylogenetic research. However, the mitochondrial (mt) genome of E. sagittatum remains uncharacterized, limiting comprehensive insights into its genomic evolution. RESULTS: In this study, we assembled the first complete mt genome of E. sagittatum employing Illumina and Nanopore sequencing technology and subsequently investigated comparative analysis with its closely related species. The mt genome of E. sagittatum was assembled as a multi-branched structure with a length of 339,191 bp, within a GC content of 46.91%. Our annotation results have shown 39 protein-coding genes (PCGs), 22 tRNA genes, three rRNA genes and four pseudogenes in the E. sagittatum mt genome. The analysis of sequence repeats has detected 79 simple sequence repeats (SSRs), 10 tandem repeats and 255 dispersed repeats in the E. sagittatum mt genome. A total of 720 C to U RNA editing sites of the 34 PCGs was predicted in E. sagittatum. The codons exhibited a strong preference for A or U bases in the E. sagittatum mt genome. The analysis of nucleotide diversity (Pi) highlighted differences in genetic variability across the tested genes, with atp9 gene exhibiting the highest genetic variation. Selection pressure analysis showed that most genes were affected by negative selection during evolution, whereas ccmB, rps10, and rps12 underwent positive selection in different plants. Additionally, a Bayesian phylogenetic tree showed that E. sagittatum was closely related to E. wushanense and E. pubescens. In total of 14 homologous fragments totaling 8,954 bp were identified between the cp and mt genomes of E. sagittatum. CONCLUSIONS: This study presents the first assembled and annotated mt genome of E. sagittatum, which provides a valuable genetic resource for the Epimedium genus and lays the foundation for investigating the phylogenetic relationship and genetic variation of this invaluable medicinal plant.

Epimedium

Assessment of Genetic Diversity and Population Structure on Azadirachta indica A. Juss. in an Urban Metropolitan: Ahmedabad, India.

Azadirachta indica (A. indica) A. Juss., commonly known as Neem, is a valuable multipurpose tree with profound medicinal properties and socioeconomic importance, widely recognized since ancient Ayurvedic times. Despite its prominence, knowledge about its genetic diversity within the metropolitan area of Ahmedabad is limited. This study marks the first in-depth exploration of the genetic diversity and population structure of A. indica in Ahmedabad. The authenticity of the species was validated through DNA barcoding, and a Geographical Information System (GIS) was used to collect the samples. A total of 35 A. indica accessions were analyzed using five Inter Simple Sequence Repeat (ISSR) primers. Genetic diversity and population structure were evaluated using Inter Simple Sequence Repeat (ISSR) markers through polymorphism assessment, clustering, ordination, and Bayesian population structure analyses. ISSRs revealed a high level of polymorphism (75.66%), indicating substantial genetic variability among accessions. An analysis of genetic diversity indices revealed low to moderate diversity (Hs = 0.14, Ht = 0.217, I = 0.217). Analysis of Molecular Variance (AMOVA) analysis depicted 81% variation within the population and 19% among the population. Low to moderate genetic differentiation (Gst = 0.319) and moderate gene flow (Nm = 1.06) indicated that urban development has not hindered gene flow among populations. Mantel's test revealed a weak but significant correlation between genetic and geographic distances, suggesting limited isolation by distance. The estimated ΔK using STRUCTURE exhibited two subpopulations, representing two gene pools for A. indica accessions (K = 2). Collectively, these patterns indicate that urbanization has not severely disrupted genetic connectivity in A. indica, reflecting its resilience and adaptive potential in a metropolitan environment. These findings provide pivotal knowledge for further understanding the genetic diversity and population structure of A. indica in one of the fastest-growing cities in India, which can be utilized for new breeding programmes, sustainable development and future conservation strategies around the globe.

India

Comparative Analysis of Chloroplast Genomes Reveals Molecular Evolution and Phylogenetic Relationships in Fraxinus (Fraxinus mandshurica).

Fraxinus mandshurica (Manchurian ash) is an ecologically and economically valuable hardwood tree native to Northeast Asia, yet its genomic resources remain limited. We assembled its complete chloroplast (cp) genome (155,559 bp) using hybrid PacBio and Illumina sequencing and performed comparative, phylogenetic, and evolutionary analyses. The cp genome exhibits a typical quadripartite structure encoding 132 gene copies, comprising 114 unique genes (80 protein-coding, 30 tRNA, and 4 rRNA genes), with 18 genes duplicated in the inverted repeat (IR) regions. Simple sequence repeat analysis revealed dominance of mononucleotide A/T repeats. Phylogenetic analysis of 53 complete cp genomes strongly supported the monophyly of Oleaceae and resolved F. mandshurica as sister to the North American F. nigra, consistent with previously proposed Miocene intercontinental dispersal scenarios between East Asia and North America. Most protein-coding genes were under strong purifying selection (Ka/Ks << 1), whereas petB, rpl2, and several ndh genes showed elevated Ka/Ks values that are suggestive of altered selective constraint but are based on very few substitutions and are therefore not, on their own, evidence of positive selection. Nucleotide diversity (Pi) analysis identified 15 hypervariable intergenic spacers (mean Pi = 0.067), among which trnM-CAU-rps14, ndhJ-ndhK, and petL-petG represent promising candidate barcode regions requiring further validation. This study provides a high-quality, fully annotated cp genome of F. mandshurica and a valuable genomic resource for future phylogenetic, population genetic, and conservation studies of this important genus.

Fraxinus

Mitogenome assembly and phylogenetic relationships of Phalaris arundinacea.

INTRODUCTION: As a perennial herb of Poaceae, Phalaris arundinacea plays key roles in grazing, production, and soil and water conservation because of its well-developed rhizomes and seed dispersal. We assembled and annotated the first mitogenome of P. arundinacea to support evolutionary and taxonomic research. METHODS: We assembled and annotated the first complete mitochondrial genome of P. arundinacea by integrating Illumina short reads with Nanopore long reads via a hybrid assembly strategy. The genome architecture was comprehensively characterized, encompassing codon usage bias, repetitive sequence organization, and inter-organellar genetic exchange with the chloroplast genome. RESULTS AND DISCUSSION: Assembly of the P. arundinacea mitogenome revealed two circular structures with a combined length of 526,717 bp. The genome comprised a set of 37 protein-coding genes (PCGs), 27 tRNAs, and 8 rRNAs, with the rRNA genes exhibiting full assembly (100% coverage). The mitochondrial genome contained 154 forward and 164 palindromic repeats, along with 25 tandem repeats and 124 simple sequence repeats (SSRs). Notably, 102 SSRs were distributed on contig1, predominantly in tetrameric form. Furthermore, 376 RNA editing sites were predicted. A total of 104 fragments were integrated into the mitochondrial genome from the chloroplast, amounting to 55,866 bp of transferred sequence. Finally, phylogenetic analysis of 28 plant mitogenomes placed P. arundinacea closest to species within the genus Poa (P. chaixii and P. pratensis). Comparative analysis of non-synonymous-to-synonymous substitution rate (Ka/Ks) ratios across divergent species revealed that the mitochondrial genome of P. arundinacea underwent stabilizing evolutionary dynamics, characterized by predominant purifying selection with several lineage-specific variations in selective pressure. Our findings support the close phylogenetic relationship between P. arundinacea and species of the genus Poa and provide a reference mitochondrial genome resource for future comparative studies within Phalaris that incorporate broader taxon sampling. These results support deeper phylogenetic investigations of P. arundinacea and facilitate future work on its germplasm characterization and applied use.

Phalaris arundinacea

The large mitochondrial genome of Syndiclis anlungensis (Lauraceae): Genome structure, comparative analysis, and phylogenetic relationships among Syndiclis species.

The complete mitochondrial genome (mitogenome) of Syndiclis anlungensis, a critically endangered tropical tree, was determined in this study. The mitogenome spans 2,368,454&#xa0;bp across four contigs and harbors 41 protein-coding genes, 22 tRNA genes, and three rRNA genes. Potential mutation regions, including 1317 repeat sequences and 698 simple sequence repeats (SSRs), were accurately located in the S. anlungensis mitogenome. Sixty-five transferred fragments of the repeats were found between its mitochondrial and chloroplast genomes. When compared to three other Laurales mitogenomes, extensive gene order shuffling is evident, leaving only five conserved gene clusters intact. Codon usage analysis reveals a pronounced A/T bias in both mitochondrial and chloroplast genes, and three mitochondrial genes (atp9, rps19, and sdh3) stand out for their high divergence across eleven Syndiclis taxa. Selection analyses indicate strong purifying pressure on rpl2, rpl16, and sdh3 (Ka/Ks&#xa0;<&#xa0;1), with no positive selection detected. Using 41 mitochondrial protein-coding gene sequences from sixteen and three individuals of Syndiclis and Beilschmiedia species, respectively, our phylogenetic tree recovers Syndiclis as monophyletic, with two well-supported clades: one includes S. anlungensis, S. chinensis, S. lotungensis, S. marlipoensis, and a putative new Syndiclis species from Yunnan; the other contains S. furfuracea, S. hongkongensis, S. kwangsiensis, and three putative new Syndiclis species from Guangdong and Vietnam.

Genome, Mitochondrial

Identifying inversions with breakpoints in the Dystrophin gene through long-read sequencing: report of two cases.

BACKGROUND: Duchenne Muscular Dystrophy (DMD) is an X-linked disorder caused by mutations in the DMD gene, with large deletions being the most common type of mutation. Inversions involving the DMD gene are a less frequent cause of the disorder, largely because they often evade detection by standard diagnostic methods such as multiplex ligation probe amplification (MLPA) and whole exome sequencing (WES). CASE PRESENTATION: Our research identified two intrachromosomal inversions involving the dystrophin gene in two unrelated families through Long-read sequencing (LRS). These variants were subsequently confirmed via Sanger sequencing. The first case involved a pericentric inversion extending from DMD intron 47 to Xq27.3. The second case featured a paracentric inversion between DMD intron 42 and Xp21.1, inherited from the mother. In both cases, simple repeat sequences (SRS) were present at the breakpoints of these inversions. CONCLUSIONS: Our findings demonstrate that LRS is an effective tool for detecting atypical mutations. The identification of SRS at the breakpoints in DMD patients enhances our understanding of the mechanisms underlying structural variations, thereby facilitating the exploration of potential treatments.

Humans

A transcriptome-wide approach for rapid pathotype discrimination of Puccinia striiformis f. sp. tritici in north-western India.

Stripe rust of wheat caused by Puccinia striiformis f. sp. tritici (Pst) remains a major constraint to wheat production in India due to the rapid evolution and frequent emergence of virulent pathotypes. Rapid and reliable discrimination of Pst pathotypes is essential for effective resistance deployment and surveillance. In the present study, transcriptome-wide simple sequence repeats (SSRs) and single nucleotide polymorphisms (SNPs) were exploited to develop and validate molecular markers for pathotype-specific detection of Pst pathotypes prevalent in North India (110S119, 238S119, 46S119, 110S84 and 78S84). Microsatellite mining from 6103 core orthologous clusters comprising 51,127 transcripts mined 14,634 SSR loci, from which 93 primer pairs were synthesized. However, only three SSR markers exhibited polymorphism indicating limited discrimination potential of expressed sequence-derived (EST) SSRs for pathotype differentiation. In contrast, SNP discovery through stringent variant calling and filtration yielded 186 pathotype-specific homokaryotic SNPs, of which 56 high-confidence loci were selected for Kompetitive Allele-Specific PCR (KASP) assay development. A total of 48 KASP markers were synthesized and 14 demonstrated clear pathotype- or cluster-specific polymorphism representing substantially higher resolution than SSR markers. The high SNP-to-KASP conversion efficiency (~&#x2009;95%) and reproducible fluorescence-based clustering emphasize the robustness of KASP assay. Comparative evaluation revealed that SNP-based KASP markers provide superior discriminatory capacity for closely related Pst pathotypes and represent a promising complementary molecular approach for rapid identification of predominant Indian Pst pathotypes. The validated marker panel developed in this study can complement conventional virulence phenotyping and field pathogenomics approaches for surveillance of currently known pathotypes, while continued refinement may accommodate future changes in pathogen populations.

India

SSR marker development for analysis of the genetic diversity and identification of species and infraspecific ranks in the genus Phyllostachys.

Bamboo plants possess important ecological, economic, and cultural values. However, it is difficult to accurately identify them on the basis of their morphological traits alone. Here, based on the whole-genome data of moso bamboo (Phyllostachys edulis) and its 20 forms, we conducted preliminary identification and comparative analyses of simple sequence repeats (SSRs) to develop molecular markers. In total, 3,835,632 SSR loci were identified from 31,537.81&#xa0;Mb of genomic sequences, among which dinucleotide SSRs were the most abundant. Most SSRs were located in intergenic regions, whereas relatively fewer were in genic regions. In addition, we found that SSR-containing genes involved in plant hormone signal transduction may be associated with the morphogenesis of moso bamboo, which was speculated to be related to differential gene expression patterns among different forms. Furthermore, 206 SSR primer pairs with polymorphisms were obtained to analyse the genetic diversity of moso bamboo and its forms, which exhibited moderate polymorphism. The proportion of genetic variation among species within the genus Phyllostachys was 58%, while that within species was 42%. Moso bamboo and its 20 forms had relatively close genetic relationships and low genetic differentiation, while 20 species of the genus Phyllostachys were clustered into three groups with distinct levels of genetic diversity. Finally, DNA fingerprints and molecular identity cards were constructed for 20 moso bamboo forms and 20 species of the genus Phyllostachys using core SSR markers. These results provide novel SSR markers for bamboo identification, germplasm conservation, and molecular marker-assisted breeding.

Microsatellite Repeats

Development and validation of whole-genome SSR markers in sugar beet (Beta vulgaris L.).

Sugar beet (Beta vulgaris L.) is an important sugar and cash crop worldwide. To systematically characterize SSR (Simple Sequence Repeat) loci across sugar beet chromosomes and enable the precise identification of germplasm resources, this study conducted a genome-wide scan for SSR loci, analyzed their distribution patterns, and determined their genotypes using resequencing data from 123 sugar beet varieties. The results revealed an abundance of SSR loci in the sugar beet genome, with a total of 135, 379 identified, from which 135, 344 pairs of SSR primers were designed (135, 344 primer pairs successfully designed; 35 loci failed to meet design criteria). Specifically, 31, 748 primer pairs were designed based on SSRs located in unassigned scaffolds, and 103, 596 primer pairs from SSRs assigned to the nine chromosomes. Through bioinformatic analysis, we identified 28, 768 SSR primers located in multi-copy genes with PIC (Polymorphism Information Content) &#x2265; 0.5, and 2, 326 SSR markers located in single-copy genes residing in various genic regions (among which 543 had PIC &#x2265; 0.5, with the highest reaching 0.776). PCR (Polymerase Chain Reaction) validation confirmed 20 robust and polymorphic markers producing clear and reproducible bands. Among them, 10 SSR primers located in multi-copy genes exhibited three or more polymorphic types, and 10 markers located in single-copy genes displayed 2-3 polymorphic types. The most polymorphic marker, YCD-4-2, detected 11 polymorphic types across 48 varieties. Furthermore, to explore markers with potential functional significance, we annotated the genes harboring SSR markers located in single-copy genes. The results showed that 1, 264 SSRs located in single-copy genes were localized to 967 genes, which are significantly enriched in pathways related to carbohydrate metabolism, stress responses, and plant-pathogen interactions. The 20 validated markers and the 2, 326 SSRs located in single-copy genes provided in this study can be directly applied to fingerprinting of sugar beet varieties, seed purity testing, and marker-assisted selection, thus representing a practical resource for molecular breeding.

genome-wide

Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing.

Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.

Journal Article

Heavy metal stress in native plant species: investigating phytoremediation potential through physiological and ISSR/SCoT molecular assessments.

In emerging countries, increased industrial activity has a significant impact on economic growth and urban development. However, the acceleration of industrial processes is accompanied by the release of contaminants such as heavy metals. According to the World Health Organization, one-fourth of all human diseases are caused by environmental contaminants, including heavy metals, which can impair numerous organs such as the neurological system, liver, and reproductive systems. This increased efforts to find effective and sustainable methods to remove heavy metals. Phytoremediation is an environmentally benign method of removing heavy metals using specific plants. Thus, from industrially contaminated locations, common native plant species of Lactuca serriola, Sisymbrium irio, Chenopodium murale, and Cynanchum acutum were selected for this study to assess the mechanisms of their molecular and physiological tolerance. Soil and plants were tested for heavy metals (Cd, Pb, and Cu), and contaminated locations were classified as low and highly polluted. Measurements were made of soluble sugar, protein, secondary metabolites, malondialdehyde, and H2O2. Additionally, inter simple sequence repeat (ISSR), start codon targeted (SCoT), and genomic template stability GTS were used. In heavily polluted areas, all plant species exhibit elevated amounts of sugar, proteins, H2O2, MDA, and secondary metabolites, while total phenolics showed a unique significant interaction (plant-location), where Cynanchum exhibited a hyper-stress phenolic accumulation to cope with toxicity, whereas Chenopodium maintained genomic stability with balanced phenolic level. Based on these findings, both Cynanchum acutum and Chenopodium murale demonstrate superior potential for phytoremediation and warrant further investigation for ecological restoration.

Heavy metal

EST-SSR based genetic polymorphism among Lablab (Lablab purpureus L. Sweet) accessions contrasting for drought stress at seedling stage.

Lablab is a multipurpose and the most drought-tolerant (DT) crop compared with its relatives. Despite its potential, Lablab is still an underutilized crop with a lack of improved varieties in many countries. The DT (D349, D147, HA4, D363, D352, D359, D348, D311, D55 and D250) and drought-susceptible (DS) (D271, D66, D106, D6, D26, D255, D28, D186, D95, and D258) accessions were earlier identified according to their morphological and biochemical responses to moisture stress at the seedling stage. These accessions were used to establish genetic polymorphism among the accessions contrasting for drought stress based on the Expressed Sequence Tag-Simple Sequence Repeats (EST-SSR) markers. The CTAB protocol was employed for the genomic DNA extraction. After DNA quality and quantity verification, the PCR was conducted using 16 EST-SSR primer pairs specific to the Lablab. The products were separated through the horizontal polyacrylamide gel electrophoresis (hPAGE). Discriminating ability of the markers and primers' efficiency were evaluated based on various genetic parameters. Principal Coordinate Analysis (PCoA) was performed to estimate the distance matrix among the population and among the accessions. While cluster analysis was processed to trace the genetic relationship among the accessions, dendrogram was constructed to decipher their genetic relationship. Analysis of Molecular Variance (AMOVA) was finally computed to quantify the diversity level and genetic relationship among the population, and among the accessions. A low polymorphism (GD&#x2009;=&#x2009;0.19) was observed between the DT and DS accessions, likely due to limited discriminatory power of the EST-SSR markers. However, the PCoA, cluster analysis and AMOVA identified DT (D147, HA4, and D349) and DS (D106, D95, and D271) accessions as strongly contrasting populations under drought stress, with D147, HA4, D349, D363, D359, D352, and D348 further recommended as DT accessions. Given the low polymorphism observed, further validation using more informative molecular markers and advanced genomic approaches is recommended to improve the identification of drought-tolerance genes and related QTLs to support Lablab breeding programs.

Expressed Sequence Tags

Molecular investigation of the progenitors, origin and domestication patterns of diploid Chinese old garden roses.

BACKGROUND AND AIMS: Chinese old garden roses are major contributors to the genetic development of modern roses. The RoKSN gene is associated with continuous flowering in roses and is proposed to have originated from Chinese wild roses. However, the wild roses that are implicated in the breeding of Chinese old garden roses and the origin of the RoKSN locus remain unidentified. We collected 25 of the most renowned and classic diploid Chinese old garden roses along with all related wild roses from East Asia. These roses were analysed with the aim of identifying the wild species that contributed to the genetic composition of Chinese old garden roses. In addition, we aimed to infer the geographical origin of the RoKSN gene and to develop a schematic overview of hybrid domestication of Chinese old garden roses. METHODS: We compared the haplotypes of internal transcribed spacers (nrITS), six nuclear single-copy genes and three chloroplast genes between Chinese old garden roses and wild roses. Additionally, we assessed genetic organization using 21 expressed sequence tag-simple sequence repeats to identify potential donor species that contributed to the emergence of these cultivars. Primers were designed for RoKSN to allow comparison of the gene across the entire distribution range of Rosa sect. Chinenses. KEY RESULTS: Our findings confirmed that the majority of rose cultivars are descendants of early hybridization events. Rosa chinensis var. spontanea, R. odorata var. gigantea and R. multiflora var. cathayensis were the primary donors for the 25 cultivar roses. Chinese old garden roses were categorized into four groups. Ten cultivars were hybrids between R. chinensis var. spontanea and R. multiflora var. cathayensis, thereby forming the 'Old Blush' group. Five cultivars were hybrids between 'Old Blush' and the R. kwangtungensis species complex, thereby forming the 'Slater's crimson' group. Six cultivars were hybrids between 'Old Blush' and R. odorata var. gigantea, thereby forming the 'Tea Rose' group, and three cultivars were hybrids that evolved from more than three donors. Moreover, we observed relatively close genetic proximity among Chinese old garden roses with an identical RoKSN-copia gene that is responsible for continuous flowering, which indicates a single origin for this retrotransposon-containing allele. Additionally, we determined that the haplotypes of the RoKSN-copia gene predominantly occurred in the Sichuan Basin region. In contrast, R. chinensis cultivated in the Ya'an region showed no markers of hybridization and displayed a genetic composition that was close to that of the wild species R. chinensis var. spontanea. This cultivar may represent the earliest mutated individual that bears the RoKSN-copia gene and may have served as a bridge from wild species to continuous-flowering old rose cultivars. CONCLUSIONS: The study provides crucial evidence that elucidates the origin of cultivated roses and lays the groundwork for further analysis of the breeding history of Chinese old garden roses using genomic data.

Domestication

Genomic signature and evolutionary history of completely cleistogamous lineages in the non-photosynthetic orchid Gastrodia.

Despite a long-standing interest since Darwin's time, the genomic implications of obligate self-fertilization remain elusive. Complete cleistogamy-the obligate production of closed, self-pollinating flowers-represents an extreme reproductive strategy. Here, we present the genomic profiles and evolutionary history of two lineages of the mycoheterotrophic orchid Gastrodia, both of which independently acquired complete cleistogamy, based on detailed sampling and a combination of simple sequence repeat (SSR), multiplexed ISSR genotyping by sequencing (MIG-seq) and RNA-seq data. Our analysis reveals clear species delimitation, with no evidence of introgression between the completely cleistogamous species and their co-occurring allogamous sisters. Intriguingly, all analyses indicate that both the completely cleistogamous Gastrodia species and their allogamous sisters exhibit genetic profiles typical of self-pollinating plants. This pattern suggests that their ancestors, probably bearing allogamous flowers, had already evolved mechanisms to mitigate the deleterious effects of selfing, potentially facilitating the emergence of complete cleistogamy through benefits such as reproductive assurance, enhanced colonization ability and species reinforcement. Meanwhile, further analyses suggest that complete cleistogamy evolved very recently (possibly within the last 1000-2000 years) in these two Gastrodia lineages. Combined with the scant evidence of complete cleistogamy outside Gastrodia, our findings imply a limited and ephemeral role for complete cleistogamy in plant speciation.

Biological Evolution

Alternative splice acceptor site in MSH4 gene is responsible for male sterility conferred by ms5 in soybean.

In soybean breeding, using the recessive male-sterile ms5 gene, derived from fast neutron mutagenesis, for recurrent selection is advantageous because of the d2 locus, which controls cotyledon color in mature seeds and can be used as a phenotypic selection marker for ms5 male sterility. However, occasional self-fertilization occurs because of the elimination of d2 linkage and instability of male sterility. Elucidating the mechanism and the gene responsible for ms5 male sterility may resolve these problems. Using fine mapping with 15 simple sequence repeat (SSR) markers, we narrowed down the candidate ms5 locus to a 54-kbp region. Bulked-DNA analysis using next-generation sequencing revealed a deletion as a candidate variation in the region. This 15-bp deletion and a nucleotide substitution were identified in intron 1 of MutS homolog (GmMSH4), which modulates chromosomal recombination in meiosis. The ms5 transcript contained a novel exon with a premature termination codon. This exon originated from an alternative splice acceptor site caused by the deletion and nucleotide substitution, disrupting gene function. Co-segregation of male sterility with five independent mutations in GmMSH4 was confirmed using progeny of mutant lines. Mutations in GmMSH4 led to biased DNA partitioning during meiosis, resulting in collapsed or enlarged pollen and suggesting that ms5 male sterility is caused by the failure of pollen formation during meiosis due to the loss of function of GmMSH4. These findings could help explain the mechanism of instability of ms5 male sterility and improve the efficiency of recurrent selection using DNA markers in soybean breeding.

Glycine max

Comparative analysis of chloroplast genomes in ten holly (Ilex) species: insights into phylogenetics and genome evolution.

In order to clarify the chloroplast genomes and structural features of ten Ilex species and provide insights into the phylogeny and genome evolution of the genus Ilex, we conducted a comparative analysis of chloroplast genomes using bioinformatics methods. The chloroplast genomes of ten Ilex species were obtained, and their structural features and variations were compared. The results indicated that all chloroplast genomes in the genus Ilex exhibit a double-stranded circular structure, with sizes ranging from 157,356 to 158,018&#xa0;bp, showing minimal differences in size. The chloroplast genomes of the ten Ilex species have a relatively conservative gene count, with a total of 134 to 135 genes, including 88 or 89 protein-coding genes, and a conserved number of 8 rRNA genes. Each chloroplast genome contains 3 to 123 SSR (Simple Sequence Repeat) sites, predominantly composed of mononucleotide and trinucleotide repeats, with no detection of pentanucleotide or hexanucleotide repeats. The variation in dispersed repeat sequences among Ilex species is minimal, with a total repeat sequence number ranging from 1 to 14, concentrated in the length range of 30 to 42 base pairs. The expansion and contraction of chloroplast genome boundaries among Ilex species are relatively stable, with only minor variations observed in individual species. Variations in non-coding regions are more pronounced than those in coding regions, with the variability in the Large Single Copy region (LSC) being the highest, while the variability in the Inverted Repeat region A (IRa) is the lowest. The divergence time among Ilex species was estimated using the MCMC-tree module, revealing the evolutionary relationships among these species, their common ancestors, and their differentiation throughout the evolutionary process. The research findings provide a valuable reference for the systematic study and molecular marker development of Ilex plants.

Genome, Chloroplast

Genomic heterozygosity and hybrid breakdown in cotton (Gossypium): different traits, different effects.

BACKGROUND: Hybrid breakdown has been well documented in various species. Relationships between genomic heterozygosity and traits-fitness have been extensively explored especially in the natural populations. But correlations between genomic heterozygosity and vegetative and reproductive traits in cotton interspecific populations have not been studied. In the current study, two reciprocal F2 populations were developed using Gossypium hirsutum cv. Emian 22 and G. barbadense acc. 3-79 as parents to study hybrid breakdown in cotton. A total of 125 simple sequence repeat (SSR) markers were used to genotype the two F2 interspecific populations. RESULTS: To guarantee mutual independence among the genotyped markers, the 125 SSR markers were checked by the linkage disequilibrium analysis. To our knowledge, this is a novel approach to evaluate the individual genomic heterozygosity. After marker checking, 83 common loci were used to assess the extent of genomic heterozygosity. Hybrid breakdown was found extensively in the two interspecific F2 populations particularly on the reproductive traits because of the infertility and the bare seeds. And then, the relationships between the genomic heterozygosity and the vegetative reproductive traits were investigated. The only relationships between hybrid breakdown and heterozygosity were observed in the (Emian22 &#xd7; 3-79) F2 population for seed index (SI) and boll number per plant (BN). The maternal cytoplasmic environment may have a significant effect on genomic heterozygosity and on correlations between heterozygosity and reproductive traits. CONCLUSIONS: A novel approach was used to evaluate genomic heterozygosity in cotton; and hybrid breakdown was observed in reproductive traits in cotton. These findings may offer new insight into hybrid breakdown in allotetraploid cotton interspecific hybrids, and may be useful for the development of interspecific hybrids for cotton genetic improvement.

Chromosomes, Plant

Strong phylogenetic signal from chloroplast genomes of three Barringtonia species provides the first genomic resources for their conservation.

BACKGROUND: The genus Barringtonia (Lecythidaceae) is a vital component of tropical coastal forests and mangrove ecosystems. Among its members, B. racemosa and B. fusicarpa are classified as Endangered and Vulnerable, respectively, due to habitat degradation and anthropogenic pressures, underscoring the urgent need for genetic studies to guide conservation. Chloroplast (cp.) genomes serve as essential resources for phylogenetic reconstruction and conservation genetics. However, the scarcity of cp. genome data for Barringtonia has limited comprehensive evolutionary and conservation-oriented investigations. RESULTS: We assembled and annotated the first complete cp. genomes of B. racemosa, B. fusicarpa, and B. acutangula. All three genomes exhibit the typical quadripartite structure, ranging from 158,959&#xa0;bp (B. racemosa) to 159,837&#xa0;bp (B. acutangula), and contain 132 genes (87 protein-coding, 37 tRNA, 8 rRNA) with a GC content of 36.68%-36.86%. Collinearity and IR boundary analyses revealed high structural conservation without large-scale rearrangements. Interspecific sequence-level variations were detected in simple sequence repeats (SSRs) and long repeats. Nucleotide diversity (&#x3c0;) analysis identified highly polymorphic regions, including rpl20 (&#x3c0;&#x2009;=&#x2009;0.080), rpoA (&#x3c0;&#x2009;=&#x2009;0.064), rps3 (&#x3c0;&#x2009;=&#x2009;0.063), and ndhF (&#x3c0;&#x2009;=&#x2009;0.060), which represent promising molecular markers for population genetics within the genus. Codon-based selection analyses (Ka/Ks) showed that all protein-coding genes are under strong purifying selection (mean Ka/Ks 0.32-0.37), with no evidence of positive selection. Pairwise genetic distances (p-distances) among Barringtonia species are extremely low (mean 0.0046), while distances to the related genus Bertholletia are ~&#x2009;6-fold higher, supporting their generic distinction. CONCLUSIONS: Phylogenetic analysis robustly supports Barringtonia as a monophyletic clade (bootstrap&#x2009;=&#x2009;100%), with B. racemosa and B. fusicarpa forming a sister lineage to B. acutangula. This study provides the first high-quality cp. genome resources for the two threatened Barringtonia species, revealing strong structural and sequence conservation but no direct chloroplast genomic correlates of endangerment. The identified polymorphic regions and repeat markers lay a foundation for future population genetics, phylogeographic studies, and conservation-oriented genetic management of these ecologically important coastal plants.

Genome, Chloroplast