PubMed HealthSearch

SEARCH · PubMed Health

Results for “non-model organisms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

17 recordsLinked to original sources

Decoding the Functional Interactome of Non-Model Organisms with PHILHARMONIC.

Despite the widespread availability of genome sequencing pipelines, many genes remain part of the genome's "dark matter," where existing inference tools cannot even begin to guess the biological function of their proteins from sequence alone. This challenge is especially pronounced in organisms that are highly evolutionarily distant from well-studied models, where homology-based methods break down. Here, we describe PHILHARMONIC, a computational method that combines deep learning-based de novo protein interaction network inference with robust unsupervised spectral clustering and remote homology to illuminate functional organization in any non-model organism. From only a sequenced proteome, we show PHILHARMONIC predicts protein functions, functional communities, and higher-order network structure with high accuracy. We validate its performance using experimental gene expression and pathway data in D. melanogaster, and we demonstrate its broad utility by analyzing temperature sensing and stress response pathways in the reef-building coral P. damicornis and its algal symbiont C. goreaui. PHILHARMONIC provides a general-purpose engine for functional discovery and biological hypothesis generation in non-model organisms, enabling systems-level insights across the full diversity of life.

Journal Article

A trainable language model with potential to modulate translation rates in non-model organisms by generating upstream untranslated region sequence libraries.

Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5'-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanistic assumptions, prior knowledge that may not be available in non-model contexts, or the screening of sequence libraries. Here, we present a simple generative approach for creating synthetic 5'-UTR libraries based solely on the genomic sequence statistics of any desired organism. The method uses a sliding-window n-gram language model applied to native 5'-UTR sequences to produce novel sequences that preserve organism-specific base distributions and motifs without hard-coding specific motifs or mechanistic rules into inflexible statistical templates. We have applied this approach to the model bacterium Escherichia coli and the non-model probiotic Limosilactobacillus reuteri. Libraries of approximately 1,000 sequences were generated for each organism, from which about 100 unique sequences were experimentally tested for translation of a fluorescent reporter protein. In both organisms, the synthetic libraries yielded a broad range of translation levels from this relatively small number of tested variants. Sequences derived from an organism's own genomic statistics provided a more uniformly distributed range of translation rates in that organism than sequences derived from the other species. Correlations of individual sequence performance across the two species were weak, and thermodynamic predictions of ribosome binding strength showed very little predictive power, especially in the non-model L. reuteri. The results demonstrate that simple statistical language model approaches applied to genomic data can generate functional translational regulatory sequence libraries without detailed mechanistic knowledge or explicit reference to consensus motifs. The approach requires minimal computational resources, avoids reproducing native sequences, and can be readily applied to any organism with a sequenced genome. This strategy may lower technical barriers to expression tuning in non-model organisms.

5' Untranslated Regions

Detection of mitochondrial tDRs in killifish embryos and other non-model organisms.

In recent years a diversity of small noncoding RNAs have been identified that originate from the mitochondrial genome. These mitosRNAs are often dominated by tRNA-derived small RNAs (mito-tDRs). Differential expression of mito-tDRs is associated with responses to stress. They also appear to be expressed differentially during development and their expression may be enriched in stress-tolerant animals. Very little is currently known about roles or modes of action of these sequences, although they are implicated in a diversity of processes such as cell cycle regulation, mRNA stability, regulation of ROS production, and import of proteins into the mitochondrion. To better understand the various roles these sequences may play, it is critical that we understand their diversity, cellular location, and the context for their expression. This protocol outlines the methodologies used to detect mitosRNAs, including mito-tDRs, in embryos and cells of the annual killifish Austrofundulus limnaeus. We highlight critical steps in the isolation of RNA, creation of sequencing libraries, bioinformatics processing of sequence data, and methods for validation of expression that support a robust discovery pipeline for mitosRNAs even from species with incomplete reference genome sequences.

Animals

MKMC enables reference-free transcriptomic analysis using k-mer representations.

Traditional RNA-seq analysis depends heavily on genome alignment and gene annotation, limiting its utility in non-model organisms and introducing biases that can obscure regulatory complexity. We present MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer-based statistics to detect biological variation without requiring alignment. MKMC integrates fast k-mer counting, abundance matrix generation, normalization, dimensionality reduction, and differential analysis into a unified workflow. Across diverse datasets, MKMC recapitulates key biological signals-including sex differences in killifish liver-and matches alignment-based pipelines in differential expression analysis and transcriptomic age prediction. Notably, MKMC detects isoform-specific events missed by traditional methods, one of which we validated using in situ hybridization. These results reveal previously hidden isoform-level regulatory events that contribute to sex- and age-associated transcriptional programs. MKMC offers a robust, extensible alternative to alignment-based approaches, enabling transcriptomic discovery across both model and non-model systems. While we focus here on RNA-seq as a primary application, MKMC is broadly applicable to any k-mer-based analysis of next-generation sequencing data.

MKMC

Protocol to predict gene expression from transcriptomic data using PREDICT.

Linking DNA sequence variation to context-specific transcriptional programs is a critical challenge in regulatory genomics, especially for non-model organisms. Here, we present PREDICT, a modular Python package for discovering cis-regulatory elements and transcription factor binding motifs. We describe steps to identify enriched k-mers from differentially expressed genes, map them to known motifs, quantify their impact on gene expression, and visualize motif co-occurrences. PREDICT provides a robust, k-mer-based approach to uncover regulatory logic in diverse genomic systems. For complete details on the use and execution of this protocol, please refer to Yen et al. and Liu et al.1,2.

Gene Expression Profiling

Simulation of CRISPR/Cas9-mediated gene editing for the Vitellogenin gene in Apis mellifera.

CRISPR/Cas9 genome editing provides a powerful framework for interrogating gene function in Apis mellifera. Yet, empirical application remains challenging due to biological constraints, including haplodiploid genetics, narrow embryonic injection window, and the social rearing requirements that complicate functional validation. These constraints necessitate in silico pre-screening to maximize editing success before resource-intensive wet-lab implementation. Within the omnigenic framework, which distinguishes core regulatory genes from peripheral loci buffered by network effects, vitellogenin (Vg) represents an optimal target which is ancestrally dedicated to yolk provisioning; it has been co-opted to orchestrate diverse non-reproductive functions including longevity, stress resistance, immunity, and social behavior. We developed a computational pipeline to design a list of 57 and 56 candidate guide RNAs (gRNA) for targeted Vg knockout, evaluating candidate sites in both functional exons 2 and 3 based on structural accessibility and frameshift efficiency. Comparative analysis revealed complementary strengths in two top-best candidates from initial target pool of predicted gRNAs. The gRNA targeting exon 2 exhibits weaker secondary structure (ΔG = -0.25 kcal/mol versus -2.10 kcal/mol for exon 3), aligning with empirical evidence that sites with ΔG > -1.0 kcal/mol achieve 2-5 × higher Cas9 binding efficiency. This site yielded moderate frameshift frequency (77.8%; 61.9 percentile). Conversely, the predicted editing outcome for the gRNA targeting exon 3, despite stronger structural constraints, demonstrated superior functional disruption metrics demonstrating very high frameshift frequency (88.3%; 95.2 percentile), high in silico editing precision, minimal microhomology-mediated repair bias, and reproducible outcomes wherein nearly all predicted indels disrupt the coding sequence. Protein structure and domain analyses further predict that frameshift edits will generate a truncated protein missing all downstream functional domains. We recommend parallel empirical validation of both exon 2 and exon 3 targets to resolve the trade-off between structural accessibility (favoring higher editing rates) and frameshift efficacy (favoring complete loss-of-function). This dual-target strategy accommodates uncertainty in in vivo performance while maximizing the probability of generating informative phenotypes. Our in silico framework enables rational CRISPR design in non-model organisms by computationally balancing biophysical accessibility with functional impact, accelerating functional genomics in species where empirical optimization faces substantial biological constraints.

Animals

GeneExt: a gene model extension tool for enhanced single-cell RNA-seq analysis.

MOTIVATION: Incomplete gene models negatively impact single-cell gene expression quantification. This is particularly true in non-model species where often gene 3' ends are inaccurately annotated, while most scRNA-seq methods only capture the 3' transcript region. This results in many genes being incorrectly quantified or not detected. RESULTS: GeneExt leverages scRNA-seq data to refine gene annotations. We exemplify GeneExt usage and its impact on the gene expression quantification of eight non-model organism single-cell atlases. By extending and homogenizing gene annotations, our tool will help improve biological interpretation and cross-species comparisons of cell type expression atlases. AVAILABILITY: GeneExt is available at https://github.com/sebepedroslab/GeneExt (DOI: https://doi.org/10.5281/zenodo.18712940) under a GNU General Public license, together with test data and usage instructions.

Software

Extensive longevity and DNA virus-driven adaptation in nearctic Myotis bats.

The genus Myotis is one of the largest clades of bats, and exhibits some of the most extreme variation in lifespans among mammals alongside unique adaptations to viral tolerance and immune defense. To study the evolution of longevity-associated traits and infectious disease, we generated cell lines and near-complete genome assemblies for 8 closely related species of Myotis. Using genome-wide screens of positive selection, analyses of structural variation, and functional experiments in primary cells, we identify new patterns of adaptation contributing to longevity, cancer resistance, and viral interactions in bats. We show that the recurrent evolution of longevity seen in Myotis leads to some of the highest predicted increases in cancer risk across mammals and demonstrate a unique DNA damage response in primary cells of the long-lived M. lucifugus. We also find evidence of abundant adaptation in response to DNA viruses - but not RNA viruses - in Myotis and other bats in sharp contrast with other mammals, potentially contributing to the role of bats as reservoirs of zoonoses. Together, our results demonstrate how genomics and primary cells derived from diverse taxa uncover the molecular bases of extreme adaptations in non-model organisms.

Aging

A century of research on the Planctomycetota bacterial phylum, previously known as Planctomycetes.

One hundred years after planctomycetes were discovered and 50 years since the first isolate was successfully cultured, this bacterial phylum remains enigmatic in many ways. In the last few decades, a significant effort to characterize new isolates has resulted in >150 described species, allowing a more comprehensive analysis of their features. However, metagenomic studies reveal that a diverse group of planctomycetes has yet to be cultured and characterized, and that many biological surprises are yet to be revealed. This is the case for the recently discovered phagotrophic Candidatus Uabimicrobium, which challenges our understanding of the distinction between prokaryotes and eukaryotes. The unique biology of planctomycete cells, such as their ability to divide without the FtsZ protein, their complex structure and characteristic morphology, their relatively large genomes containing many genes with unknown function, and their variable metabolic capabilities, imposes significant barriers for researchers. Although ubiquitous, the precise ecological roles of planctomycetes in various environments are still not fully understood. However, their distinctive metabolism opens the door to a large number of potential biotechnological applications, which are beginning to be unveiled. In this article, we first review the historical milestones in planctomycetes research and describe the pioneers of the field. We then describe the controversies and their resolutions, we highlight the past discoveries and current interrogations related to planctomycetes, and discuss the ongoing challenges that hinder a comprehensive understanding of their biology. We end up with directions for exploring the biology and ecological roles of these fascinating organisms.

Bacteria

Induced Pluripotent Stem Cells in Non-Model Species: Applications and Challenges.

Induced pluripotent stem cells have revolutionized biomedical research-yet the vast majority of life on Earth remains beyond their reach. Non-model species lack the annotated genomes, validated reagents, and species-specific culture infrastructure that make iPSC technology routine in humans and mice, and this infrastructure deficit, compounded by genuine biological differences in pluripotency network architecture across taxa, is what has kept the field narrow. The deep conservation of the core pluripotency network across vertebrates suggests that reprogramming may, in principle, be achievable across a far broader range of species than currently demonstrated-though the extent to which this holds across more divergent taxa remains to be established. This review consolidates current progress and future potential of iPSC technology across five domains: technical reprogramming challenges and advances; conservation applications including genetic rescue, in vitro gametogenesis, and de-extinction; medical applications within a one medicine framework; agricultural applications spanning disease resistance, climate resilience, and cultured meat; and species-specific iPSC-derived systems in ecotoxicology. Throughout, we distinguish what has been demonstrated from what remains aspirational and identify the priorities that will determine whether the iPSC revolution can be extended-rigorously and at scale-beyond model organism research.

Induced Pluripotent Stem Cells

Genome-wide characterization of NOD-like receptor genes links NLR repertoire evolution to spleen immune responses after Aeromonas hydrophila challenge in the Chinese spiny frog (Quasipaa spinosa).

NOD-like receptors (NLRs) are cytosolic pattern-recognition receptors that detect pathogen-associated and damage-associated molecular patterns and mediate innate immune signaling in vertebrates. However, the genomic repertoire, evolutionary diversification, and infection-associated expression of NLR genes remain poorly defined in non-model amphibians. In this study, 66 NLR genes were identified from the Chinese spiny frog (Quasipaa spinosa) genome and designated as QsNLR1-QsNLR66. These genes were unevenly distributed across chromosomes and were classified into three phylogenetic groups, with most members exhibiting conserved motif architectures. Gene duplication analysis indicated that dispersed duplication was the main contributor to QsNLR expansion. Synteny analysis detected five conserved orthologous gene pairs between Q. spinosa and Pelophylax nigromaculatus, suggesting partial conservation of NLR genomic organization between the two amphibians. Ka/Ks analysis showed that several duplicated gene pairs, including NLRC3-like/QsNLR36 and NLRC3-like/QsNLR50, exhibited Ka/Ks ratios greater than one, suggesting potential sequence divergence after duplication. Spleen RNA sequencing (RNA-seq) after Aeromonas hydrophila challenge revealed enrichment of immune-related Gene Ontology (GO) terms and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways. Weighted gene co-expression network analysis linked several QsNLRs to infection-associated modules, among which QsNLR57 was co-expressed with CYBB, ADAM17, SPI1, and HK2. RT-qPCR using time-matched phosphate-buffered saline (PBS) controls showed distinct temporal patterns, with stronger induction of QsNLR29, QsNLR57, and QsNLR66 and weaker or delayed responses of QsNLR50 and QsNLR56. These results characterize the NLR repertoire of Q. spinosa and identify infection-associated QsNLR candidates for future studies of antibacterial immunity in amphibians.

Animals

Expansion of satellite DNAs derived from transposable elements in beetles with reduced diploid numbers.

Repetitive DNA sequences are ubiquitous in eukaryotic genomes, significantly influencing their structure, function, and evolution. They can facilitate genomic rearrangements, contributing to chromosomal and genomic diversity. Chrysomelidae (Coleoptera) beetles are known for their highly diverse karyotypes and heterochromatin distribution. In this study, we advanced the understanding of the intricate relationship between satellite DNA-like sequences (named here solely as satDNA) and genome organization/reshuffling using three species of Eumolpinae chrysomelids. We investigated the satellitomes of three species with divergent karyotypes that had undergone independent chromosomal fusions: Colaspis laeta (2n = 22, Xyp), with a conserved karyotype; Endocephalus bigatus (2n = 10, neo-XY); and Iphimeis dives (2n = 14, neo-XY). Our comparative analysis revealed highly divergent patterns of satDNA origin, organization, and evolution. In species with reduced chromosome numbers and neo-sex chromosomes, we observed a high abundance of transposable element-related (TE-related) satDNAs. In Colaspis laeta, the sex chromosomes (Xyp) showed an advanced level of differentiation. However, in the species with a reduction in diploid number, such a level of differential enrichment of repetitive DNAs was not observed in the sex chromosomes, indicating an early stage of differentiation. Our findings support the hypothesis that chromosomal rearrangements and reorganization of repetitive DNA sequences are connected, with extensive reshuffling observed in species with reduced diploid numbers. Moreover, the data reinforce the involvement of TEs in satDNA origin, which could spread widely throughout the genome, including euchromatic areas. This study provides new insights into the evolutionary dynamics of repetitive DNAs in non-model species, emphasizing the impact of chromosomal rearrangements on genome architecture and evolution.

Animals

Integrative proteomics and bioinformatics pipelines for PTM profiling.

Post-translational modifications (PTMs) regulate protein function across all life forms and allow plants to respond rapidly to biotic and abiotic stress. Over 450 PTM types have been described across organisms, of which 23-33 have been experimentally confirmed in plants, including phosphorylation, acetylation, methylation, glycosylation, ubiquitination, and sumoylation. These modifications are highly dynamic and often reversible, and frequently act in combination, or "crosstalk," to fine-tune cellular processes. Advances in high-resolution mass spectrometry and large-scale genome sequencing continue to expand the catalogue of known PTM sites, while machine learning and deep learning approaches increasingly support prediction of PTM site localization and function. Unlike broader surveys of plant PTMs, this review focuses specifically on O-phosphorylation and Lys-N(ε)-acetylation, the two best-characterized and most extensively crosstalking PTMs in plants, and integrates four perspectives: the historical development of proteomic and bioinformatics approaches to these modifications; current mass spectrometry-based workflows and enrichment strategies; the bioinformatics tools and databases available for their analysis; and the technical and species-related challenges, particularly in non-model plants, that currently limit their study. We close by outlining priority directions for future research, including multi-omics integration, AI-based prediction, and the translation of PTM knowledge into crop stress resilience and breeding applications.

Protein Processing, Post-Translational

Sex-Chromosome-Dependent Ageing in Female Heterogametic Methylomes.

Recent research in humans and both model and non-model animals has shown that DNA methylation (DNAm), an epigenetic modification, is one of the mechanisms underlying the ageing process. DNAm-based indices predict mortality and provide valuable insights into biological ageing mechanisms. Although sex-dependent differences in lifespan are ubiquitous and sex chromosomes are thought to play an important role in sex-specific ageing, they have been largely ignored in epigenetic ageing studies. We characterised the genome-wide distribution of age-related CpG (Cytosine-phosphate-Guanine) sites from longitudinal samples in two avian species (zebra finch and jackdaw), including for the first time the avian sex chromosomes (Z and the female-specific, haploid W). In both species, we find a small fraction of the CpG sites to show age-related changes in DNAm with the majority of them being located on the haploid, female-specific W chromosome, where DNAm levels predominantly decrease with age. Age-related CpG sites were over-represented on the zebra finch but under-represented on the jackdaw Z chromosome. Our results highlight distinct age-related changes in sex chromosome DNAm compared to the rest of the genome in two avian species, suggesting this previously understudied feature of sex chromosomes may be instrumental in sex-dependent ageing. Moreover, studying the DNAm of sex chromosomes might be particularly useful in ageing research, facilitating the identification of shared (sex-dependent) age-related pathways and processes between phylogenetically diverse organisms.

Animals

Comparative Genomics of Sex-Determination-Related Genes Reveals Shared Evolutionary Patterns Between Bivalves and Mammals, but Not Fruit Flies.

The molecular basis of sex determination (SD), while being extensively studied in model organisms, remains poorly understood in many animal groups. Bivalves, a diverse class of molluscs with a variety of reproductive modes, represent an ideal yet challenging clade for investigating SD and the evolution of sexual systems. However, the absence of a comprehensive framework has limited progress in this field, particularly regarding the study of sex-determination-related genes (SRGs). In this study, we performed a genome-wide sequence evolutionary analysis of the Dmrt, Sox and Fox gene families in more than 40 bivalve species. For the first time, we provide an extensive and phylogenetically aware dataset of these SRGs, and we find support for the hypothesis that Dmrt-1L and Sox-H may act as primary sex-determining genes by showing their high levels of sequence diversity within the bivalve genomic context. To validate our findings, we studied the same gene families in two well-characterised systems, mammals and fruit flies (genus Drosophila). In the former, we found that the male sex-determining gene Sry exhibits a pattern of amino acid sequence diversity similar to that of Dmrt-1L and Sox-H in bivalves, consistent with its role as master SD regulator. In contrast, no such pattern was observed among genes of the fruit fly SD cascade, which is controlled by a chromosomic mechanism. Overall, our findings highlight similarities in the sequence evolution of some mammal and bivalve SRGs, possibly driven by a comparable architecture of SD cascades. This work underscores once again the importance of employing a comparative approach when investigating understudied and non-model systems.

Animals

Satellite DNAs in Drosophila koepferae (repleta group) reveal patterns of origin, chromosomal organization, transcription, and turnover in the buzzatii cluster.

Satellite DNAs (satDNAs) are non-coding tandem repeats that can comprise more than 20% of eukaryotic genomes. They contribute to structural and regulatory processes in the genome and often evolve rapidly, shaping early stages of genetic differentiation between populations and species. Although Drosophila has long served as a model for studying satDNA biology, little is known about satDNAs in non-model Drosophila species, particularly within the repleta group, one of the most species-rich lineages in the genus. To reduce such bias, several studies have focused on the buzzatii cluster (repleta group). However, D. koepferae remained the only species lacking comprehensive satDNA data, limiting comparative analyses. Here, we used publicly available genomic sequencing data from two D. koepferae populations (Argentina and Bolivia) to characterize their satDNA content. Both populations share the same set of five satDNAs (CDSTR8, CDSTR138, CDSTR230, DBC-150 and CDSTR177), which together account for ~ 0,9% of the genomic DNA. We show that CDSTR177 originated through amplification of an internal segment of the Galileo transposable element, an event restricted to D. koepferae. All satDNAs localize to heterochromatic regions, with CDSTR138 most likely associated to the centromeres of most chromosomes. Transcripts from all satDNAs were detected, although at low levels. Our results provide new insights into the origin, genomic contribution, expression and evolution of satDNAs in the buzzatii cluster, support incipient differentiation between Argentinean and Bolivian populations of D. koepferae and contribute to clarifying the phylogenetic position of this species within the buzzatii cluster.

Animals

FANTASIA suite: a reproducible and configurable framework for embedding-based functional annotation of proteins.

Embedding-based annotation transfer is increasingly used for protein function inference due to protein language models capture sequence, structural, and functional signals that may extend beyond conventional pairwise similarity. However, systematic application of these approaches requires control over model choice, reference composition, lookup parameters, evidence traceability, and output formats. We developed the FANTASIA suite, a configurable framework for embedding-based functional annotation of proteins. The suite combines a database-backed implementation for reproducible and extensible analyses with a portable flat-file implementation for rapid local annotation and pipeline integration. Using non-model and model-organism proteomes, we show that larger neighbourhood sizes remain practical for proteome-scale analyses and that taxonomy and sequence-identity filtering support leakage-aware benchmarking. We also compare the supported models with baseline methods through external CAFA5 evaluation and provide practical guidance based on empirical evidence variables. FANTASIA provides a controlled, scalable, and reproducible framework for extending functional annotation across the rapidly expanding diversity of sequenced organisms.

Software