PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genomics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Evolution of paralogous genes: Reconstruction of genome rearrangements through comparison of multiple genomes within Staphylococcus aureus.

Analysis of evolution of paralogous genes in a genome is central to our understanding of genome evolution. Comparison of closely related bacterial genomes, which has provided clues as to how genome sequences evolve under natural conditions, would help in such an analysis. With species Staphylococcus aureus, whole-genome sequences have been decoded for seven strains. We compared their DNA sequences to detect large genome polymorphisms and to deduce mechanisms of genome rearrangements that have formed each of them. We first compared strains N315 and Mu50, which make one of the most closely related strain pairs, at the single-nucleotide resolution to catalogue all the middle-sized (more than 10 bp) to large genome polymorphisms such as indels and substitutions. These polymorphisms include two paralogous gene sets, one in a tandem paralogue gene cluster for toxins in a genomic island and the other in a ribosomal RNA operon. We also focused on two other tandem paralogue gene clusters and type I restriction-modification (RM) genes on the genomic islands. Then we reconstructed rearrangement events responsible for these polymorphisms, in the paralogous genes and the others, with reference to the other five genomes. For the tandem paralogue gene clusters, we were able to infer sequences for homologous recombination generating the change in the repeat number. These sequences were conserved among the repeated paralogous units likely because of their functional importance. The sequence specificity (S) subunit of type I RM systems showed recombination, likely at the homology of a conserved region, between the two variable regions for sequence specificity. We also noticed novel alleles in the ribosomal RNA operons and suggested a role for illegitimate recombination in their formation. These results revealed importance of recombination involving long conserved sequence in the evolution of paralogous genes in the genome.

Amino Acid Sequence↗

Detection of alien chromosomes from S-genome species in the addition/substitution lines of bread wheat and visualization of A-, B- and D-genomes by GISH.

A modified approach based on the GISH technique for detecting introgressed chromosomes/chromosome arms from closely related S-genome species to wheat genome and for visualization of A-, B- and D-genomes of Triticum aestivum L. (genome AABBDD, 2n = 6x = 42) is presented. For detecting alien chromosomes we investigated two lines of bread wheat, one is an addition line with a pair of chromosome No. 4 short arms from Aegilops searsii (4SsS) and a wheat substitution line with a pair of chromosomes No. 6 from Ae. longissima (6S1). A hybridization mixture consists of two differently labelled DNAs, one from the line used for chromosome spread preparations, and the second from origin species of alien chromosomes. The latter adds different color in the regions of its hybridization showing the presence of alien chromosomes by creating a strong and easily detected combined signal. For discriminating A-, B-, and D-genome chromosomes, the hybridization mixture of differently labelled total DNA from Ae. tauschii--the proposed progenitor of D-genome (detected red) and T. dicoccoides (genome AABB) (detected green) were used. The high temperature of hybridization allows high precision annealing of chromosome/probe sequences and at the same time it sharpens differences between reassociation kinetics of eu- and heterochromatin revealing chromosome substructure. A pre-annealing step increases probe specificity. As a result, we observed brown chromosomes of A-genome, banded green chromosomes of B-genome and red chromosomes of D-genome. Inter genomic invasion of the sequences from A/B-genomes to D-genome has been detected.

Chromosomes↗

Dynamic evolution of genomes and the concept of genome space.

A new era in the elucidation of genome evolution has been heralded with the availability of numerous genome sequences. With these data, it has been possible to study evolutionary processes at a greater level of detail in order to characterize features such as gene shuffling, genome rearrangements, base bias composition, and horizontal gene transfer. In this paper, we discuss the evolutionary implications of significant rearrangements within genomes as well as characteristic genomic regions that have been conserved across genomes. This is based on our analysis of orthologous and paralogous genes. We argue that genome plasticity has most likely contributed substantially to the dynamic evolution of genomes. We also describe the characteristic mosaic features of an archaea genome that is comprised of both bacterial and eukaryal elements. Here we investigate base compositional differences as well as the similarity of this species' genes to either bacteria or eukarya. We conclude that these features can be largely explained by the mechanism of horizontal gene transfer. Finally, we introduce the concept of genome space which is defined as the entire set of genomes of all living organisms. We explain its usefulness to describe as well as to gain deeper insight into the general features of the dynamic genomic evolutionary process.

Archaea↗

Identification of genomic species in Agrobacterium biovar 1 by AFLP genomic markers.

Biovar 1 of the genus Agrobacterium consists of at least nine genomic species that have not yet received accepted species names. However, rapid identification of these organisms in various biotopes is needed to elucidate crown gall epidemiology, as well as Agrobacterium ecology. For this purpose, the AFLP methodology provides rapid and unambiguous determination of the genomic species status of agrobacteria, as confirmed by additional DNA-DNA hybridizations. The AFLP method has been proven to be reliable and to eliminate the need for DNA-DNA hybridization. In addition, AFLP fragments common to all members of the three major genomic species of agrobacteria, genomic species G1 (reference strain, strain TT111), G4 (reference strain, strain B6, the type strain of Agrobacterium tumefaciens), and G8 (reference strain, strain C58), have been identified, and these fragments facilitate analysis and show the applicability of the method. The maximal infraspecies current genome mispairing (CGM) value found for the biovar 1 taxon is 10.8%, while the smallest CGM value found for pairs of genomic species is 15.2%. This emphasizes the gap in the distribution of genome divergence values upon which the genomic species definition is based. The three main genomic species of agrobacteria in biovar 1 displayed high infraspecies current genome mispairing values (9 to 9.7%). The common fragments of a genomic species are thus likely "species-specific" markers tagging the core genomes of the species.

Bacterial Typing Techniques↗

Identification of probable genomic packaging signal sequence from SARS-CoV genome by bioinformatics analysis.

AIM: To predict the probable genomic packaging signal of SARS-CoV by bioinformatics analysis. The derived packaging signal may be used to design antisense RNA and RNA interfere (RNAi) drugs treating SARS. METHODS: Based on the studies about the genomic packaging signals of MHV and BCoV, especially the information about primary and secondary structures, the putative genomic packaging signal of SARS-CoV were analyzed by using bioinformatic tools. Multi-alignment for the genomic sequences was performed among SARS-CoV, MHV, BCoV, PEDV and HCoV 229E. Secondary structures of RNA sequences were also predicted for the identification of the possible genomic packaging signals. Meanwhile, the N and M proteins of all five viruses were analyzed to study the evolutionary relationship with genomic packaging signals. RESULTS: The putative genomic packaging signal of SARS-CoV locates at the 3' end of ORF1b near that of MHV and BCoV, where is the most variable region of this gene. The RNA secondary structure of SARS-CoV genomic packaging signal is very similar to that of MHV and BCoV. The same result was also obtained in studying the genomic packaging signals of PEDV and HCoV 229E. Further more, the genomic sequence multi-alignment indicated that the locations of packaging signals of SARS-CoV, PEDV, and HCoV overlaped each other. It seems that the mutation rate of packaging signal sequences is much higher than the N protein, while only subtle variations for the M protein. CONCLUSIONS: The probable genomic packaging signal of SARS-CoV is analogous to that of MHV and BCoV, with the corresponding secondary RNA structure locating at the similar region of ORF1b. The positions where genomic packaging signals exist have suffered rounds of mutations, which may influence the primary structures of the N and M proteins consequently.

Amino Acid Sequence↗

CREAT: A CRISPR-Based Genome Trimming Strategy for Systematic Identification of Dispensable Regions and Rapid Genome Reduction.

The construction of minimal-genome microbes offers an ideal platform for understanding fundamental biological processes and synthetic biology, yet the research is hindered by incomplete lists of essential genes in microbes and by multiple rounds of genome trimming with a trial-and-error nature. To address this, we introduce CREAT (CRISPR-based genome trimming with a multi-homology-arm template)-a streamlined approach that integrates CRISPR-targeted genome cleavage and homology arm walking to classify essential from non-essential genomic subregions, thus providing the basis for predicting essential genes in a given organism. These essential genes were then assembled into synthetic gene cassettes for one-step replacement of the targeted non-deletable genomic regions for further genome trimming. Eight consecutive rounds of CREAT genome trimming achieved a 20.8% reduction in genome size in Saccharolobus islandicus. Furthermore, Cas9-based CREAT genome trimming was developed for Bacillus subtilis and Escherichia coli, with efficiency greatly enhanced by the λ-Red recombinase in the latter. Together, this iterative application of CREAT provides a scalable and generally applicable strategy for rapidly constructing minimal genomes across diverse microorganisms.

CRISPR-Cas Systems↗

Genomics, genetic epidemiology, and genomic medicine.

Medical science is on the threshold of unparalleled progress as a result of the advent of genomics and related disciplines. Human genomics, the study of structure, function, and interactions of all genes in the human genome, promises to improve the diagnosis, treatment, and prevention of disease. This opportunity is the result of the recent completion of the Human Genome Project. It is anticipated that genomics will bring to physicians a powerful means to discover hereditary elements that interact with environmental factors leading to disease. However, the expected transformation toward genomics-based medicine will occur over decades. It will require efforts of many scientists and physicians to begin now to sort out the vast amounts of information in the human genome and translate it to meaningful applications in clinical practice. Meanwhile, practicing physicians and health professionals need to be trained in the principles, applications, and limitations of genomics and genomic medicine. Only then will we be in a position to benefit patients, which is the ultimate goal of accelerating scientific progress in medicine. In this inaugural article, we introduce and discuss concepts, facts, and methods of genomics and genetic epidemiology that will be drawn on in the forthcoming topics of the clinical genomics series.

Chromosome Mapping↗

Selection of genomic sequences that bind tightly to Ff gene 5 protein: primer-free genomic SELEX.

Single-stranded DNA or RNA libraries used in SELEX experiments usually include primer-annealing sequences for PCR amplification. In genomic SELEX, these fixed sequences may form base pairs with the central genomic fragments and interfere with the binding of target molecules to the genomic sequences. In this study, a method has been developed to circumvent these artificial effects. Primer-annealing sequences are removed from the genomic library before selection with the target protein and are then regenerated to allow amplification of the selected genomic fragments. A key step in the regeneration of primer-annealing sequences is to employ thermal cycles of hybridization-extension, using the sequences from unselected pools as templates. The genomic library was derived from the bacteriophage fd, and the gene 5 protein (g5p) from the phage was used as a target protein. After four rounds of primer-free genomic SELEX, most cloned sequences overlapped at a segment within gene 6 of the viral genome. This sequence segment was pyrimidine-rich and contained no stable secondary structures. Compared with a neighboring genomic fragment, a representative sequence from the family of selected sequences had about 23-fold higher g5p-binding affinity. Results from primer-free genomic SELEX were compared with the results from two other genomic SELEX protocols.

Base Sequence↗

Genome size and the accumulation of simple sequence repeats: implications of new data from genome sequencing projects.

The relationship between the level of repetitiveness in genomic sequences and genome size has been re-investigated making use of the rapidly growing database of complete eubacterial and archaeal genome sequences combined with the fragmentary but now large amount of data from eukaryotic genomes. Relative simplicity factors (RSFs), which measure the repetitiveness of sequences, were calculated and significantly simple motifs (SSMs), which identify the kinds of sequences that are repeated, were identified. A previously reported correlation between genome size and repetitiveness was confirmed, but it was shown that the higher RSFs seen in eukaryotic genomes also reflect a generally higher level of repetitiveness independent of genome size differences. Differences in genome size are responsible for about 10% of the variance in RSF seen between species. The spectrum of SSMs seen within a genome differed markedly within the eubacteria but less so in eukaryotes and, particularly, in archaea. Species with SSM spectra that differ from the norm tend also to have high RSFs for their genome size and to be pathogens that make use of repetitive sequences to avoid host defence responses. Some of the variance in repetitiveness seen in other species may therefore also reflect the action of selection, although other forces such as variation in the effectiveness of mechanisms for regulating slippage errors of replication, may also be important.

Animals↗

Yeast genomic databases and the challenge of the post-genomic era.

Since the completion of the yeast genome sequence in 1996, three genomic databases, the Saccharomyces Genome Database, the Yeast Proteome Database, and MIPS (produced by the Munich Information Center for Protein Sequences), have organized published knowledge of yeast genes and proteins onto the framework of the genome. Now, post-genomic technologies are producing large-scale datasets of many types, and these pose new challenges for knowledge integration. This review first examines the structure and content of the three genomic databases, and then draws from them and other resources to examine the ways knowledge from the literature, genome, and post-genomic experiments is stored, integrated, and disseminated. To better understand the impact of post-genomic technologies, 20 collections of post-genomic data were analyzed relative to a set of 243 previously uncharacterized genes. The results indicate that post-genomic technologies are providing rich new information for nearly all yeast genes, but data from these experiments is scattered across many Web sites and the results from these experiments are poorly integrated with other forms of yeast knowledge. Goals for the next generation of databases are set forth which could lead to better access to yeast knowledge for yeast researchers and the entire scientific community.

Databases, Genetic↗

Genomes for Nurses: Understanding and Overcoming Barriers to Nurses Utilizing Genomics.

Background: Genomic testing is an increasingly important technology within pediatric oncology that aids in cancer diagnosis, provides prognostic information, identifies therapeutic targets, and reveals underlying cancer predisposition. However, nurses lack basic knowledge of genomics and have limited self-assurance in using genomic information in their daily practice. This single-institution project was carried out at an academic pediatric cancer hospital in the United States with the aim to explore the barriers to achieving genomics literacy for pediatric oncology nurses. Method: This project assessed barriers to genomic education and preferences for receiving genomics education among pediatric oncology nurses, nurse practitioners, and physician assistants. An electronic survey with demographic questions and 15 genetics-focused questions was developed. The final survey instrument consisted of nine sections and was pilot-tested prior to administration. Data were analyzed using a ranking strategy, and five focus groups were conducted to capture more-nuanced information. The focus group sessions lasted 40 min to 1 hour and were recorded and transcribed. Results: Over 50% of respondents were uncomfortable with or felt unprepared to answer questions from patients and/or family members about genomics. This unease ranked as the top barrier to using genomic information in clinical practice. Discussion: These results reveal that most nurses require additional education to facilitate an understanding of genomics. This project lays the foundation to guide the development of a pediatric cancer genomics curriculum, which will enable the incorporation of genomics into nursing practice.

Humans↗

i-Genome: a database to summarize oligonucleotide data in genomes.

BACKGROUND: Information on the occurrence of sequence features in genomes is crucial to comparative genomics, evolutionary analysis, the analyses of regulatory sequences and the quantitative evaluation of sequences. Computing the frequencies and the occurrences of a pattern in complete genomes is time-consuming. RESULTS: The proposed database provides information about sequence features generated by exhaustively computing the sequences of the complete genome. The repetitive elements in the eukaryotic genomes, such as LINEs, SINEs, Alu and LTR, are obtained from Repbase. The database supports various complete genomes including human, yeast, worm, and 128 microbial genomes. CONCLUSIONS: This investigation presents and implements an efficiently computational approach to accumulate the occurrences of the oligonucleotides or patterns in complete genomes. A database is established to maintain the information of the sequence features, including the distributions of oligonucleotide, the gene distribution, the distribution of repetitive elements in genomes and the occurrences of the oligonucleotides. The database can provide more effective and efficient way to access the repetitive features in genomes.

Alu Elements↗

An acquisition account of genomic islands based on genome signature comparisons.

BACKGROUND: Recent analyses of prokaryotic genome sequences have demonstrated the important force horizontal gene transfer constitutes in genome evolution. Horizontally acquired sequences are detectable by, among others, their dinucleotide composition (genome signature) dissimilarity with the host genome. Genomic islands (GIs) comprise important and interesting horizontally transferred sequences, but information about acquisition events or relatedness between GIs is scarce. In Vibrio vulnificus CMCP6, 10 and 11 GIs have previously been identified in the sequenced chromosomes I and II, respectively. We assessed the compositional similarity and putative acquisition account of these GIs using the genome signature. For this analysis we developed a new algorithm, available as a web application. RESULTS: Of 21 GIs, VvI-1 and VvI-10 of chromosome I have similar genome signatures, and while artificially divided due to a linear annotation, they are adjacent on the circular chromosome and therefore comprise one GI. Similarly, GIs VvI-3 and VvI-4 of chromosome I together with the region between these two islands are compositionally similar, suggesting that they form one GI (making a total of 19 GIs in chromosome I + chromosome II). Cluster analysis assigned the 19 GIs to 11 different branches above our conservative threshold. This suggests a limited number of compositionally similar donors or intragenomic dispersion of ancestral acquisitions. Furthermore, 2 GIs of chromosome II cluster with chromosome I, while none of the 19 GIs group with chromosome II, suggesting an unidirectional dispersal of large anomalous gene clusters from chromosome I to chromosome II. CONCLUSION: From the results, we infer 10 compositionally dissimilar donors for 19 GIs in the V. vulnificus CMCP6 genome, including chromosome I donating to chromosome II. This suggests multiple transfer events from individual donor types or from donors with similar genome signatures. Applied to other prokaryotes, this approach may elucidate the acquisition account in their genome sequences, and facilitate donor identification of GIs.

Algorithms↗

Genome analysis of multiple pathogenic isolates of Streptococcus agalactiae: implications for the microbial "pan-genome".

The development of efficient and inexpensive genome sequencing methods has revolutionized the study of human bacterial pathogens and improved vaccine design. Unfortunately, the sequence of a single genome does not reflect how genetic variability drives pathogenesis within a bacterial species and also limits genome-wide screens for vaccine candidates or for antimicrobial targets. We have generated the genomic sequence of six strains representing the five major disease-causing serotypes of Streptococcus agalactiae, the main cause of neonatal infection in humans. Analysis of these genomes and those available in databases showed that the S. agalactiae species can be described by a pan-genome consisting of a core genome shared by all isolates, accounting for approximately 80% of any single genome, plus a dispensable genome consisting of partially shared and strain-specific genes. Mathematical extrapolation of the data suggests that the gene reservoir available for inclusion in the S. agalactiae pan-genome is vast and that unique genes will continue to be identified even after sequencing hundreds of genomes.

Amino Acid Sequence↗

Sequence analysis of the genome of the unicellular cyanobacterium Synechocystis sp. strain PCC6803. II. Sequence determination of the entire genome and assignment of potential protein-coding regions.

The sequence determination of the entire genome of the Synechocystis sp. strain PCC6803 was completed. The total length of the genome finally confirmed was 3,573,470 bp, including the previously reported sequence of 1,003,450 bp from map position 64% to 92% of the genome. The entire sequence was assembled from the sequences of the physical map-based contigs of cosmid clones and of lambda clones and long PCR products which were used for gap-filling. The accuracy of the sequence was guaranteed by analysis of both strands of DNA through the entire genome. The authenticity of the assembled sequence was supported by restriction analysis of long PCR products, which were directly amplified from the genomic DNA using the assembled sequence data. To predict the potential protein-coding regions, analysis of open reading frames (ORFs), analysis by the GeneMark program and similarity search to databases were performed. As a result, a total of 3,168 potential protein genes were assigned on the genome, in which 145 (4.6%) were identical to reported genes and 1,257 (39.6%) and 340 (10.8%) showed similarity to reported and hypothetical genes, respectively. The remaining 1,426 (45.0%) had no apparent similarity to any genes in databases. Among the potential protein genes assigned, 128 were related to the genes participating in photosynthetic reactions. The sum of the sequences coding for potential protein genes occupies 87% of the genome length. By adding rRNA and tRNA genes, therefore, the genome has a very compact arrangement of protein- and RNA-coding regions. A notable feature on the gene organization of the genome was that 99 ORFs, which showed similarity to transposase genes and could be classified into 6 groups, were found spread all over the genome, and at least 26 of them appeared to remain intact. The result implies that rearrangement of the genome occurred frequently during and after establishment of this species.

Bacterial Proteins↗

Systemic polyomavirus genome increase and dissemination of capsid-defective genomes in mammary gland tumor-bearing mice.

BALB/c mice that developed tumors 7 to 8 months following neonatal infection by polyomavirus (PYV) wild-type strain A2 were characterized with respect to the abundance and integrity of the viral genome in the tumors and in 12 nontumorous organs. These patterns were compared to those found in tumor-free mice infected in parallel. Six mice were analyzed in detail including four sibling females with mammary gland tumors. In four of five mammary gland tumors, the viral genome had undergone a unique deletion and/or rearrangement. Three tumor-resident genomes with an apparently intact large T coding region were present in abundant levels in an unintegrated state. Two of these had undergone deletions and rearrangements involving the capsid genes and therefore lacked the capacity to produce live virus. In the comparative organ survey, the tumors harboring replication-competent genomes contained by far the highest levels of genomes of any tissue. However, the levels of PYV genomes in other organs were elevated by up to 1 to 2 orders of magnitude compared to those detected in the same organs of tumor-free mice. The genomes found in the nontumorous organs had the same rearrangements as the genomes residing in the tumors. The original wild-type genome was detected at low levels in a few organs, particularly in the kidneys. The data indicate that a systemic increase in the level of viral genomes occurred in conjunction with the induction of tumors by PYV. The results suggest two novel hypotheses: (i) that genomes may spread from the tumors to the usual PYV target tissues and (ii) that this dissemination may take place in the absence of capsids, providing an important path for a virus to escape from the immune response. This situation may offer a useful model for the spread of HPV accompanying HPV-induced oncogenesis.

Animals↗

Genomic inbreeding coefficients and inbreeding depression of semen production traits at genome-wide and chromosomal levels in Japanese Holstein bulls.

We aimed to estimate inbreeding coefficients and the effects of inbreeding depression on semen production traits at both the genome-wide and chromosomal levels. We utilized pedigree data for 19,921 animals, single nucleotide polymorphism (SNP) data on 5700 Japanese Holstein bulls, and 52,193 semen collection records from 775 bulls. We estimated 4 different inbreeding coefficients, namely a pedigree-based coefficient (FPED) and 3 genomic coefficients derived from SNP data. The genomic coefficients consisted of one based on the genomic relationship matrix (FGRM), one based on runs of homozygosity (ROH), and one based on homozygous-by-descent (HBD) segments (FHBD). These genomic coefficients were estimated at both the genome-wide and chromosomal levels. Furthermore, we investigated the effects of these coefficients on semen production traits: semen volume (VOL), sperm concentration (CON), sperm number (NUM), and sperm motility (MOT). In the genome-wide-level analysis, inbreeding coefficients increased markedly in bulls born after 2009, coinciding with the introduction of genomic selection. Significant inbreeding depression of VOL was found. At the chromosomal level, the inbreeding coefficients for most chromosomes showed a similar trend to the genome-wide metrics, although some (e.g., chr10 and chr20) exhibited a more pronounced trend. Suggestive inbreeding effects were detected on specific chromosomes for all traits (chr1 and chr22 for VOL, chr24 and chr29 for CON, chr1, chr12, and chr27 for NUM, chr10 and chr18 for MOT), including the traits that were not significant at the genome-wide level. Our results highlight that chromosomal-level analysis provides information complementary to whole-genome metrics, offering a more detailed perspective for managing inbreeding effects. To mitigate the adverse effects of inbreeding on semen production traits, future breeding programs would benefit from the control of inbreeding effects on high-risk chromosomal regions.

Genomic inbreeding coefficient↗

Comparative genomics and phylogenetic analysis of three Malvaceae species on the basis of chloroplast genomes.

INTRODUCTION: The Malvaceae family shows rich species diversity and has substantial economic and medicinal value. However, the frequent interspecific hybridization among members of this family has resulted in confused phylogenetic relationships among the groups, limiting the usefulness of traditional classification methods. METHODS: This study aimed to investigate the phylogenetic relationships among selected taxa of Malvaceae by evaluating 23 chloroplast (CP) genomes, including three newly assembled CP genomes. Among these three genomes, the CP genome of Hibiscus schizopetalus L. was reported for the first time, while the CP genomes of Alcea rosea L. and Hibiscus grewiifolius L., which have been deposited in NCBI, were re-analyzed here alongside newly generated data for comparative purposes. In addition, 20 downloaded CP genomes encompassing 13 genera were analyzed using SNPs in whole CP genomes data. RESULTS: The results showed that the genomes ranged from 160,403 to 161,978 base pairs in length and consisted of small single copies (SSCs) and large single copies (LSCs) separated by two inverted repeat sequences (IRs), forming a typical quadripartite circular structure. The entire genome sequence showed relative conservation across species in terms of structure, GC content, codon usage, and gene composition. The mutation sites were mainly located in the LSC and SSC regions, and the variability in the non-coding regions was higher than that in the coding regions. The nucleotide polymorphism (Pi) analysis identified the non-coding regions such as ndhF-rpl32 and psbZ-trnG as high variable hotspots. A maximum likelihood phylogenetic tree was constructed based on SNPs in whole CP genomes data. The phylogenetic analysis divided these 23 species into five highly supported clades. It also revealed a close sister-group relationship between Abelmoschus and Hibiscus species, suggesting that Hibiscus may have a separate lineage from okra species. DISCUSSION: In conclusion, the increasing availability of CP genome resources will enhance our understanding of the classification and evolutionary patterns of the Malvaceae family. The development of molecular markers will provide important molecular evidence for precise identification and classification revision of plants in this family.

Malvaceae↗