PubMed HealthSearch

SEARCH · PubMed Health

Results for “genetic code expansion”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Accessing isotopically labeled proteins containing genetically encoded phosphoserine for NMR with optimized expression conditions.

Phosphoserine (pSer) sites are primarily located within disordered protein regions, making it difficult to experimentally ascertain their effects on protein structure and function. Therefore, the production of 15N- (and 13C)-labeled proteins with site-specifically encoded pSer for NMR studies is essential to uncover molecular mechanisms of protein regulation by phosphorylation. While genetic code expansion technologies for the translational installation of pSer in Escherichia coli are well established and offer a powerful strategy to produce site-specifically phosphorylated proteins, methodologies to adapt them to minimal or isotope-enriched media have not been described. This shortcoming exists because pSer genetic code expansion expression hosts require the genomic ΔserB mutation, which increases pSer bioavailability but also imposes serine auxotrophy, preventing growth in minimal media used for isotopic labeling of recombinant proteins. Here, by testing different media supplements, we restored normal BL21(DE3) ΔserB growth in labeling media but subsequently observed an increase of phosphatase activity and mis-incorporation not typically seen in standard rich media. After rounds of optimization and adaption of a high-density culture protocol, we were able to obtain ≥10 mg/L homogenously labeled, phosphorylated superfolder GFP. To demonstrate the utility of this method, we also produced the intrinsically disordered serine/arginine-rich region of the SARS-CoV-2 Nucleocapsid protein labeled with 15N and pSer at the key site S188 and observed the resulting peak shift due to phosphorylation by 2D and 3D heteronuclear single quantum correlation analyses. We propose this cost-effective methodology will pave the way for more routine access to pSer-enriched proteins for 2D and 3D NMR analyses.

Humans

High throughput screening of eukaryotic release factor 1 variants to enhance noncanonical amino acid incorporation.

Noncanonical amino acids (ncAAs) enable diversification of protein functions, but the efficiency of genetic code expansion (GCE) in eukaryotes is hindered by competition between suppressor tRNAs and release factors. Prior work has identified eukaryotic release factor 1 (eRF1) mutants that improve ncAA incorporation, suggesting that screens for improved variants may lead to further enhancements. Here, we developed a high-throughput system to screen eRF1 mutants in Saccharomyces cerevisiae where eRF1 mutants are coexpressed on a plasmid alongside genomically encoded, wild-type eRF1. This strategy enabled recovery of live cells expressing eRF1 variants that enhance ncAA incorporation, even with mutants known to severely affect cell viability in the absence of WT eRF1 expression. We prepared and screened a million-member library of randomly mutated eRF1 variants for clones exhibiting improved ncAA integration phenotypes. Deep sequencing revealed a diverse set of enriched mutations across all three major domains of eRF1. Interestingly, several enriched mutations identified here are also found in naturally occurring eRF1 homologs from species that recode canonical stop codons. When eRF1 variants were combined with yeast knockout strains also known to enhance ncAA incorporation, this resulted in further improvements to efficiency, highlighting the complementarity of release factor engineering to other GCE enhancement strategies. This work demonstrates that high-throughput engineering of the eukaryotic translational apparatus is a powerful approach to identify previously unknown solutions for enhancing ncAA incorporation, with implications for elucidating and precisely manipulating the molecular functions of essential translational machinery.

Noncanonical amino acids

Genetic Incorporation of a Thioxanthone-Containing Amino Acid for the Design of Artificial Photoenzymes.

Genetically encodable photosensitizers allow the design of artificial photoenzymes to expand the scope of abiological reactions. Herein, we report the genetic incorporation of a thioxanthone-containing amino acid into a protein scaffold via an engineered pyrrolysyl-tRNA/pyrrolysyl-tRNA synthetase pair. The designer enzyme was engineered to catalyze a dearomative [2+2] cycloaddition reaction in high yields (up to>99 % yield) with excellent enantioselectivity (up to 98 : 2 e.r.). This work provides a robust and facile method for photoenzyme design and lays the foundation for the development of further photoenzymatic reactions.

Xanthones

A Unified Mechanism of +1 Ribosomal Frameshifting.

Ribosomes decode 3-nucleotide codons and move in 1-codon increments to maintain the messenger RNA (mRNA) frame thereby accurately producing the encoded protein. In special cases, including viral genomes and regulatory cellular proteins, frameshifting occurs to expand the coding repertoire of an mRNA to make more than one protein. How these frameshifting events are induced and regulated is an active area of research. Here, we discuss recent progress in the understanding of +1 frameshifting (+1FS), during which the ribosome shifts by 1 mRNA nucleotide in the 3' direction. Structural and biochemical studies yielded insights into +1FS induced by mRNA slippery sequences and transfer RNA (tRNA) stem-loop expansion or modifications. tRNAs with an additional anticodon nucleotide are explored as a biotechnology tool for expanding the genetic code in an approach termed quadruplet decoding. We revisit the challenges of the quadruplet decoding model, discuss +1FS scenarios in bacteria and eukaryotes, and propose a unifying structural mechanism for +1FS.

Frameshifting, Ribosomal

Regulatory Evolution and the Genetic Basis of Human Brain Expansion.

The evolution of the human brain is characterized by profound changes in structure and function, despite relatively limited divergence in protein-coding genes compared to other primates. This paradox has led to increasing recognition of gene regulatory elements (GREs) as primary drivers of evolutionary innovation. In this review, we synthesize current knowledge on the role of conserved noncoding elements (CNEs), human accelerated regions (HARs), and transposable element (TE)-derived sequences in shaping gene regulatory networks (GRNs) underlying brain development. Comparative analyses across humans and closely related primates, including the chimpanzee, gorilla, and orangutan, reveal that while core regulatory architectures are highly conserved, subtle changes in regulatory elements drive species-specific gene expression patterns. We highlight how CNEs provide a stable regulatory framework, whereas HARs and TE-derived elements introduce lineage-specific modifications that fine-tune neurodevelopmental processes. Advances in functional genomics, including CRISPR-based perturbations, massively parallel reporter assays, and single-cell multi-omics, have enabled direct interrogation of regulatory function, linking sequence variation to cellular phenotypes. Furthermore, we discuss how regulatory evolution contributes to both cognitive innovation and susceptibility to neurological disorders. Despite significant progress, challenges remain in establishing causal relationships between regulatory variation and phenotypic outcomes. Future integration of multi-omics data and comparative models will be essential for resolving these complexities. Together, this review provides a comprehensive framework for understanding the molecular basis of primate brain evolution through the lens of gene regulation.

Brain evolution

The proteomic origin of the genetic code.

INTRODUCTION: The origin and evolution of the genetic code is a central problem in molecular biology. Classical models have emphasized stereochemistry, frozen accidents, or adaptive optimization, often treating proteins as passive products of preexisting codes. More recent views instead portray the code as a dynamic, coevolving system shaped by reciprocal interactions among amino acids, RNA, and early catalysts. AREAS COVERED: Here, I review efforts of phylogeny reconstruction of the history of tRNA, protein structural domains, and dipeptide sequences in proteomes. These complementary approaches allow exploration of the entry of amino acids and codons into the code, and the transition from an operational RNA code in the tRNA acceptor arm to the canonical code in the anticodon loop. Evidence for ancestral synthetase enzymes with dual functions in aminoacylation and peptide-bond formation, as well as early bidirectional (sense-antisense) coding reflected in dipeptide-antidipeptide emergence is also discussed. EXPERT OPINION: The genetic code is best viewed as a proteome-driven, evolvable system in which early peptides actively shaped coding rules by stabilizing structure, expanding chemical diversity, and enhancing catalysis. This perspective connects origin-of-life studies with modern efforts of code expansion, translational engineering, and peptide-based therapeutics, highlighting the impact of the code's proteomic origin.

Genetic Code

The complete sequence of the silkworm W chromosome uncovers its rapid evolution by large-scale duplications/deletions and translocation of W-linked genes.

The complete sequence of the W chromosome, which carries feminization activity in the silkworm, is crucial for understanding the sex-determination system in Lepidoptera. However, extensive accumulation of transposons due to lack of recombination, the very rare protein-coding genes and almost no information about molecular markers has hindered full W sequencing. We report the first complete silkworm W sequence (T2T_W, 11683305 bp) obtained by combining sequencing-assembly technologies and newly developed error detection methods, evaluated with genetically mapped W-RAPD markers, W-mutants, and W-derived BAC clones. The T2T_W sequence showed that the W is composed of a massive 92% accumulation of transposons and repeat sequences, among which the main constituents are intact LTR/LINE retrotransposons indicating recent expansions. In addition to Fem clusters producing Fem piRNA (Feminizer-derived PIWI-interacting RNA), we found 26 protein-coding genes in the W sequence. These include four gene pairs encoding zinc-finger motifs designated z1:z20 and a gene encoding serine/arginine repetitive matrix protein 1-like (SRRM1-like). To identify candidate genes for female sex-determination and differentiation we also sequenced the shortest W (3.8 Mb) from a translocation mutant with feminizing activity, which harbored four conventional genes: a Fem cluster, a pair of z1:z20 isoforms, z20-S, and a SRRM1-like gene. Phylogenetic analysis revealed that z1:z20 originated from a copy of an autosomal zinc-finger gene pair, z2:z21, translocated onto the W around 2.43 Mya and subsequently amplified to yield 4 W-linked zinc-finger gene pairs. The complete W sequence revealed that large-scale deletions and amplifications played a significant role in W chromosome evolution.

Animals

3D epigenome of glial cell types in developing human cortex.

The human cortex is complex and heterogeneous, undergoing extensive expansion during development1,2. Our prior study of neurogenesis, including radial glia (RG), intermediate progenitor cells, excitatory neurons and interneurons demonstrated that chromatin looping underlies transcriptional regulation for lineage-specific genes, shedding light on how non-coding genetic variants contribute to neuropsychiatric disorders by means of cell-type-specific gene regulation3. RG have a crucial role in generating cellular diversity through both neurogenesis and gliogenesis and can be further classified into ventricular RG (vRG) and outer RG (oRG)4,5. Given their significance in cortical development, we conducted a comprehensive three-dimensional (3D) epigenomic analysis of four main glial populations, including vRG, oRG, oligodendrocyte precursor cells and microglia, from the mid-gestational human neocortex. By integrating gene expression, chromatin accessibility, DNA methylation and 3D chromatin interactions, we identified cell-type-specific candidate cis-regulatory elements (cCREs) and validated their regulatory function using transgenic mouse embryos. Using machine learning, we prioritized 112 schizophrenia risk variants within glia cCREs and further confirmed the predicted vRG enhancer disruption by the rs4449074 risk allele in vivo. Finally, oRG cCREs are enriched for human accelerated regions compared with other cCREs and a subset of human accelerated regions show activity differences from their chimpanzee orthologues that interact with genes involved in neuronal development. Our findings advance the understanding of human-specific gene regulation during corticogenesis.

Journal Article

A human-specific non-coding RNA for EFHC1, an epilepsy-associated gene, regulates neural stem cell proliferation for cortical development.

Epilepsy is a prevalent brain disorder in humans but rarely occurs naturally in other species, highlighting the potential for human-specific mechanisms in its pathogenesis, and thus, current animal models fail to recapitulate human symptoms. Comparing RNA sequencing (RNA-seq) datasets from human and mouse neural stem cells (NSCs), we identified EFHC1, a juvenile myoclonic epilepsy gene, as exhibiting a human-biased expression. EFHC1 knockdown reduced human NSC proliferation, while its overexpression in mouse embryonic brains increased cortical NSC number. Mechanistically, EFHC1 prevented endoplasmic reticulum stress, thereby reducing inflammatory activation of p38 MAPK and promoting continuous proliferation of human NSCs. We also identified pancEFHC1, a bidirectional promoter-associated non-coding RNA (pancRNA), located at the human EFHC1 promoter. Knockdown of pancEFHC1 in human NSCs increased DNA methylation to reduce EFHC1 expression, with the resulting phenotype rescued by EFHC1 overexpression. We propose that the evolutionary acquisition of pancEFHC1 has introduced a complex regulatory mechanism for EFHC1 expression that allows distinguishing it in humans.

Humans

Unraveling the genomic blueprint of the Indian black soldier fly: From genome assembly to evolutionary insights.

The black soldier fly (BSF) (Hermetia illucens) has been renowned for its sustainable bioconversion capabilities, resulting in smart protein production with wide applications in animal feed, bioenergy, and biofertilizer. However, the genetic mechanisms underlying efficient bioconversion and productivity remain poorly understood. To advance strain-specific applications and strengthen genetic resource availability, we present the whole genome sequencing (WGS) data for an Indian isolate of black soldier fly. The assembled genome was 1.46 Gb with a scaffold N50 of 172.7 Mb, and a GC content of 42.6%. Furthermore, 64.17% of genomic sequences were masked as repeated, and 14,317 protein-coding sequences were identified. Variant analysis against the reference genome identified 34.44 million variants (∼33.25 million SNPs and ∼ 1.18 million INDELs), with the majority (99.3%) classified as MODIFIER, 0.54% as LOW impact, 0.14% as MODERATE, and only 0.003% as HIGH impact. Comparative genomic analysis with other related species revealed expansions of gene families in BSF associated with Immune effector (Antimicrobial peptides (AMPs), Lysozymes, and Peptidoglycan Recognition Protein (PGRP) and Detoxification (cytochrome P450 enzymes). Notably, AMPs in the Indian isolate showed enhanced copy number variation in defensin (27) and PGRP (40) compared to reference BSF, suggesting potential regional adaptations to pathogen exposure. Collectively, this genomic data provides an improved resource for evolutionary studies, functional genomics, and targeted genetic improvement of BSF for sustainable bioconversion applications.

Comparative genomics

Establishment of four induced pluripotent stem cell lines (IGIBi028-A, IGIBi029-A, IGIBi030-A, and IGIBi031-A) from peripheral blood derived cells of Spinocerebellar ataxia Type 12 patients.

Spinocerebellar ataxia type 12 (SCA12) is a progressive late-onset neurodegenerative disorder caused by expansion of ≥ 43 trinucleotide CAG repeats in the upstream non-coding region of the PPP2R2B gene at locus 5q32 (SCA12; OMIM#604326). Clinically SCA12 patients predominately present hand tremor, gait ataxia, tremulous voice and other neurological and psychiatric features. Neuroimaging reveals degenerative changes in the cerebral cortex and cerebellum, however, the underlying disease mechanism at molecular level is still incompletely understood. Here we report generation of four induced pluripotent stem cells (iPSCs) of SCA12 patients. The established lines were positive for PPP2R2B-CAG expansion mutation and showed expression of undifferentiated hPSC state markers, three germ layer differentiation potential, normal genetic integrity and contamination-free culture.

Humans

Genetic Deletion of Cis-Regulatory Elements to Dissect the Function of the Non-coding Genome in human Preimplantation Models.

Cis-regulatory elements coordinate gene expression in a spatially and temporally controlled manner and contribute to the establishment of distinct cellular states during development. A substantial proportion of transcriptionally active cis-regulatory elements in primate embryos originated from ancient retroviral integrations into the germline. These endogenous retroviruses, also known as long terminal repeat retrotransposons, retain intrinsic regulatory activity and are often species-specific, making them strong candidates for regulating species-divergent aspects of embryonic development. Ethical and legal restrictions on human embryo research have historically limited direct investigation of gene regulation during human embryogenesis. Human naive pluripotent stem cells and three-dimensional stem cell-based blastocyst models provide alternative systems for studying early developmental processes. This protocol describes the CRISPR-Cas9-mediated deletion of endogenous retrovirus-derived cis-regulatory elements in human naive pluripotent stem cells. Preassembled Cas9 and single-guide RNA ribonucleoprotein complexes are delivered by nucleofection, followed by single-cell cloning, PCR-based genotyping, Sanger sequencing, expansion, cryopreservation, and genomic stability assessment of the edited lines. The resulting wild-type, heterozygous, and homozygous or hemizygous deletion clones provide a platform for investigating the contribution of individual endogenous retrovirus-derived elements to gene regulation in human preimplantation models. This method enables direct functional interrogation of species-specific non-coding regulatory sequences and supports the study of transcriptional mechanisms involved in early human development.

Humans

PhyClone: accurate Bayesian reconstruction of cancer phylogenies from bulk sequencing.

MOTIVATION: Cancer is driven by somatic mutations that result in the expansion of genomically distinct sub-populations of cells called clones. Identifying the clonal composition of tumours and understanding the evolutionary relationships between clones is a crucial task in cancer genomics. Bulk DNA sequencing is commonly used for studying the clonal composition of tumours, but it is challenging to infer the genetic relationship between different clones due to the mixture of different cell populations. RESULTS: In this work, we introduce a new probabilistic model called PhyClone that can infer clonal phylogenies from bulk-sequencing data. We demonstrate the performance of PhyClone on simulated and real-world datasets and show that it outperforms previous methods in terms of accuracy and sample scalability. AVAILABILITY AND IMPLEMENTATION: Source code is available on Github at: https://github.com/Roth-Lab/PhyClone under the GPL v3.0 license.

Neoplasms

Are we Prepared? Genetic Counseling for Stillbirth in the Sequencing Era.

Stillbirth affects approximately 1 in 175 pregnancies annually in the United States. Although the American College of Obstetricians and Gynecologists recommends genetic testing as part of the stillbirth evaluation, families often face barriers to obtaining a complete evaluation. Expansion of the diagnostic evaluation of stillbirth is expected to include exome/genome sequencing, with preliminary studies demonstrating its diagnostic utility. Consequently, genetic counselors (GCs) are expected to play an expanding role in post-stillbirth care. This study explored current genetic counseling practices for stillbirth and GCs' preparedness to support patients in this setting. A cross-sectional survey was distributed across four channels. Eligible participants included GCs in the United States and Canada with at least 1 year of prenatal experience. The survey assessed GC frequency and timing in stillbirth counseling, genetic testing practices, comfort addressing psychosocial needs, and perceived barriers to care. Responses were analyzed using descriptive statistics. Group comparisons were performed using Chi-square and Fisher's exact tests. Open-ended responses were coded for themes. Seventy-one responses were analyzed. Approximately half of respondents (49.3%, n = 36) reported "never/very rarely/rarely" counseling patients postpartum, despite this being the optimal time to offer genetic testing. Delivering providers (46.5%, n = 33) were often responsible for informing patients about testing and obtaining consent, compared to GCs (11.3%, n = 8). Although chromosomal microarray (CMA) is recommended as the standard of care (SOC), 12.7% (n = 9) of GCs reported not offering CMA for anomalous and non-anomalous stillbirths. Perceived barriers to SOC testing included reported lack of obstetrician awareness (91.5%, n = 65) and challenges coordinating specimen collection (90.1%, n = 64). These findings highlight barriers to SOC genetic evaluation and underscore the need to strengthen institutional protocols, enhance provider education, and develop stillbirth-specific genetic counseling guidelines. GC involvement in these efforts will be essential to promoting equitable access to comprehensive post-stillbirth care as sequencing becomes integrated into practice.

Humans

From genes to trajectories: mapping genetic influences on Huntington's disease progression.

MOTIVATION: There are many diseases with established genetic factors, such as Huntington's disease (HD), that are characterized by variable rates of progression. However, beyond the contribution of the known genetic factors - in this case the Huntingtin (HTT) gene - the impact of the full human genome on the natural progression of such diseases throughout a patient's life remains largely unknown. The increased availability of genome wide association (GWA) data in HD gene expansion carriers (HDGECs), combined with the clinical assessment scores on the same set of patients, has provided a perfect opportunity to assess the potentially broader genetic impact on the natural progression of HD. RESULTS: We present a genetics-driven, probabilistic disease progression model designed to identify and investigate the ways in which a range of genetic factors affect the natural progression of HD. When applied to a clinico-genomic HD dataset, our model identified several single nucleotide polymorphisms (SNPs) with previously unreported effects on disease progression that act at distinct stages and with varying magnitudes. This discovery may shed light on the potential mechanistic impact of previously unidentified genes on HD that may have implications for clinical management. As increasing amounts of GWA data become available more generally, we anticipate that this modeling framework will be broadly applicable to other diseases with strong genetic components. AVAILABILITY AND IMPLEMENTATION: The source code for IHDPM is available at https://github.com/BiomedSciAI/IHDPM.

Huntington Disease

Genome-wide insights into the evolutionary and demographic history of the red alga Mazzaella laminarioides: Evidence for speciation with ancient migration along the southeast Pacific coast.

The mechanisms driving lineage divergence in red algae remain unexplored, despite the group's remarkable diversity and ancient evolutionary history. The red alga Mazzaella laminarioides, a Chilean intertidal species complex composed of three parapatric cryptic lineages (North, Center, South), offers a valuable system to evaluate these processes, as its life history combines severe dispersal limitation with a haploid-diploid cycle that may influence the emergence of reproductive barriers. We reconstructed its evolutionary history using whole-genome sequencing and nuclear genome assembly of representative individuals from each lineage. Phylogenomic analyses based on 1,507 single-copy orthologs recovered three deeply divergent lineages with limited nuclear discordance consistent with incomplete lineage sorting. For both splits, demographic modelling was most consistent with an Ancient Migration scenario, although support over strict isolation was moderate, suggesting that divergence may have begun with low asymmetric ancestral gene flow followed by subsequent loss of connectivity, demographic bottlenecks, and later population expansion. Coding sequence analyses revealed lineage-specific dN/dS heterogeneity; only one South-lineage locus passed FDR correction (metaxin-1, mitochondrial protein import), with two further South-lineage candidates in chlorophyll and heme biosynthesis falling below the FDR threshold. Together, these signals suggest that divergent selective pressures on energy acquisition may have contributed to divergence at the southern end of the distribution. These results add to the small but growing body of whole-genome data for red algae and, alongside recent macroalgal studies, suggest that ancestral connectivity could be a recurrent feature of lineage divergence even in marine organisms with extremely restricted dispersal.

Rhodophyta

Mitochondrial DNA control-region and coding-region data highlight geographically structured diversity and post-domestication population dynamics in worldwide donkeys.

Donkeys (Equus asinus) have been used extensively in agriculture and transportations since their domestication, ca. 5000-7000 years ago, but the increased mechanization of the last century has largely spoiled their role as burden animals, particularly in developed countries. Consequently, donkey breeds and population sizes have been declining for decades, and the diversity contributed by autochthonous gene pools has been eroded. Here, we examined coding-region data extracted from 164 complete mitogenomes and 1392 donkey mitochondrial DNA (mtDNA) control-region sequences to (i) assess worldwide diversity, (ii) evaluate geographical patterns of variation, and (iii) provide a new nomenclature of mtDNA haplogroups. The topology of the Maximum Parsimony tree confirmed the two previously identified major clades, i.e. Clades 1 and 2, but also highlighted the occurrence of a deep-diverging lineage within Clade 2 that left a marginal trace in modern donkeys. Thanks to the identification of stable and highly diagnostic coding-region mutational motifs, the two lineages were renamed as haplogroup A and haplogroup B, respectively, to harmonize clade nomenclature with the standard currently adopted for other livestock species. Control-region diversity and population expansion metrics varied considerably between geographical areas but confirmed North-eastern Africa as the likely domestication center. The patterns of geographical distribution of variation analyzed through phylogenetic networks and AMOVA confirmed the co-occurrence of both haplogroups in all sampled populations, while differences at the regional level point to the joint effects of demography, past human migrations and trade following the spread of donkeys out of the domestication center. Despite the strong decline that donkey populations have undergone for decades in many areas of the world, the sizeable mtDNA variability we scored, and the possible identification of a new early radiating lineage further stress the need for an extensive and large-scale characterization of donkey nuclear genome diversity to identify hotspots of variation and aid the conservation of local breeds worldwide.

Animals

Complexity of schistosome vector bulinine snails in Kenya: Insights from nuclear genome size variation, complete mitochondrial genome sequence, and morphometric analysis.

Investigations of nuclear genome size, complete mitochondrial genome (mitogenome) sequence, and morphometrics were conducted on specimens of Bulinus snails (Gastropoda: Planorbidae) collected from 14 locations across the east coast, central Kenya, and western Kenya around the Lake Victoria region (November 2013 and January 2024). Flow cytometry measurements of DNA content (C-value) revealed unexpected variation in nuclear genome size, with diploid Bulinus africanus and B. forskalii species groups showing C-values ranging from 0.76 to 1.98 pg, while tetraploid B. truncatus had a C-value of 1.82 pg. Additionally, C-values for six B. globosus specimens from different localities ranged from 1.43 to 1.98 pg. These findings suggest that bulinine snails, particularly the B. africanus species group, have undergone genome expansion, whole genome duplication (polyploidization), or both, which have not been previously recognized. Next-generation sequencing was performed to determine and annotate 14 complete mitogenome sequences. Despite the well-conserved arrangement of protein-coding genes, two versions of mtDNA genome structure, distinguished by the tRNA-D (Asp) location, were found, designated as DCF (Asp-Cys-Phe) type (in the B. forskalii group and the B. truncatus/tropicus complex) and CF (Cys-Phe) type (in the B. africanus group). Phylogenetic analyses based on complete mtDNA sequences of bulinines from Kenya, along with cytochrome c oxidase subunit I (COX1) sequences from various localities across Africa, contributed to resolving species identities and provided further support for the presence of multiple or cryptic species in the taxon B. globosus. A landmark-based morphometric analysis was ineffective in distinguishing these species. This study reveals unexpected nuclear genome size variation, provides new mitogenome sequences, and highlights the limitations of morphological analysis. It offers valuable insights into the cytogenetics, polyploidy, genomics, taxonomy, and evolution of bulinines, which serve as intermediate hosts for schistosomes responsible for human urogenital schistosomiasis and intestinal schistosomiasis in domestic and wild mammals.

Animals