PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genes”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Co-expression in tissue-specific gene networks links genes in cancer-susceptibility loci to known somatic driver genes.

BACKGROUND: The genetic background of cancer remains complex and challenging to integrate. Many somatic mutations within genes are known to cause and drive cancer, while genome-wide association studies (GWAS) of cancer have revealed many germline risk factors associated with cancer. However, the overlap between known somatic driver genes and positional candidate genes from GWAS loci is surprisingly small. We hypothesised that genes from multiple independent cancer GWAS loci should show tissue-specific co-regulation patterns that converge on cancer-specific driver genes. RESULTS: We studied recent well-powered GWAS of breast, prostate, colorectal and skin cancer by estimating co-expression between genes and subsequently prioritising genes that show significant co-expression with genes mapping within susceptibility loci from cancer GWAS. We observed that the prioritised genes were strongly enriched for cancer drivers defined by COSMIC, IntOGen and Dietlein et al. The enrichment of known cancer driver genes was most significant when using co-expression networks derived from non-cancer samples of the relevant tissue of origin. CONCLUSION: We show how genes within risk loci identified by cancer GWAS can be linked to known cancer driver genes through tissue-specific co-expression networks. This provides an important explanation for why seemingly unrelated sets of genes that harbour either germline risk factors or somatic mutations can eventually cause the same type of disease.

Humans

Identification and Classification of Expressed Orphan Genes, Spurious Orphan Genes, and Conserved Genes in the Human Gut Microbiome.

Orphan genes (OGs)-genes lacking detectable homologs outside a species-are widespread in microbial genomes and are thought to contribute to their adaptation and molecular innovation. However, not all predicted OGs may represent novel functional coding sequences. False positive OGs, also called spurious OGs, can arise from gene prediction errors. We reason that OGs lacking detectable expression are more likely to be spurious. To test this, we combined large-scale metatranscriptomic profiling of the human gut microbiome with machine learning to distinguish expressed OGs from spurious ones and compare them with conserved genes (CGs) found in multiple species. Using nearly 5,000 metatranscriptome libraries, we identified ∼218,000 OGs supported by expression evidence, while ∼330,000 predicted OGs lacked detectable expression and were classified as spurious. We extracted 154 features for sequence, structural, and evolutionary properties for each gene and trained XGBoost classifiers while accounting for genomic representation. The models achieved an area under the receiver operating characteristic curve (AUC) of 0.82 in distinguishing expressed OGs from spurious OGs and an AUC of 0.93 in distinguishing expressed OGs from CGs. Interpretation based on SHAP (SHapley Additive exPlanations) revealed clear biological signals. Particularly, expressed orphans were present in more genomes than spurious ones, and expressed OGs were shorter than CGs. This work improves OG discovery and suggests that expressed OGs differ systematically from CGs and spurious OGs in sequence composition, structural constraints, and evolutionary signals.

Humans

Exploring prognostic genes in the immune microenvironment of acute myeloid leukemia via weighted gene co-expression network analysis.

BACKGROUND: Acute myeloid leukemia (AML) is a heterogeneous blood cancer that arises from transformed myeloid precursor cells in a compromised bone marrow microenvironment. This environment is essential for AML initiation, progression, and relapse. Alongside oncogenic changes in hematopoietic cells, immunological dysregulation also contributes to leukemogenesis. The present study is aimed to identify prognostic genes in stromal and immune cells associated with AML using the weighted gene co-expression network analysis (WGCNA). METHODS: Gene expression profiles were retrieved from The Cancer Genome Atlas database, and immune and stromal cell scores were calculated using the ESTIMATE (Estimation of STromal and Immune cells in MAlignant Tumor tissues using Expression data) method. These scores helped identify differentially expressed genes (DEGs), which were then used to create gene clusters through WGCNA. To explore the functions of genes linked to AML subtypes, Gene Ontology and Kyoto Encyclopedia of Genes and Genomes enrichment analyses were performed. A protein-protein interaction network was developed to identify hub genes. The top 18 hub genes were identified using the cytoHubba plug-in in Cytoscape software, and survival analysis was conducted with the Gene Expression Profiling Interactive Analysis 2 online tool. RESULTS: A total of 1097 DEGs were identified, with 601 being upregulated and 496 downregulated. WGCNA analysis indicated that the gray module, comprising 165 genes, had the strongest association with AML subtypes (Cor&#x2005;>&#x2005;0.3; P&#x2005;<&#x2005;.05). Gene Ontology enrichment analysis demonstrated that the 18 identified hub genes were predominantly associated with neutrophil activation, immune response, secretory granule membrane, and pattern recognition receptor activity. Kyoto Encyclopedia of Genes and Genomes pathway enrichment analysis revealed that the DEGs were mainly involved in pathways related to phagosome, lysosome, tuberculosis, leishmaniasis, and neutrophil extracellular trap formation. Kaplan-Meier survival analysis of the top 18 hub genes indicated that ITGAM, IL10, and CD163 were significantly correlated with survival outcomes in AML. CONCLUSION: Key stromal and immune-related genes influencing AML patient outcomes were identified, highlighting their potential as therapeutic targets. These discoveries provide deeper insights into the molecular mechanisms driving AML pathogenesis and subtype differentiation.

Leukemia, Myeloid, Acute

A sequence-based classifier distinguishes phenotype-associated genes from other gene models in plants.

Only a small fraction of annotated plant genes possess experimentally validated associations with specific phenotypes. Phenotype-associated genes have distinct structural, molecular, and evolutionary characteristics compared with nonvalidated gene models. Here, we develop a simple classifier that uses sequence and evolutionary features, which can be generated for any species with an annotated reference genome assembly, to accurately distinguish phenotype-associated genes from both the overall population of annotated gene models and a specific set of genes identified as being tolerant of premature stop mutations. A model trained solely on genes from maize (Zea mays) identifies and prioritizes rice (Oryza sativa) and Arabidopsis (Arabidopsis thaliana) genes that are highly enriched in genes with experimentally validated links to phenotypes in both of these evolutionarily distant species. Gene models predicted to have a higher probability of being linked to phenotypes display patterns consistent with known biological properties of phenotype-associated genes. Notably, the sets of genes predicted to have a high probability of being linked to phenotype variation do not consist exclusively of well-characterized gene families but included many uncharacterized gene families carrying domains of unknown function. The quantitative scores generated by this model offer a valuable resource for prioritizing and exploring the vast number of uncharacterized gene models in plants, reducing the risk of failure in future reverse genetic efforts and potentially accelerating gene discovery and functional annotation in crops.

Phenotype

Diverse evolutionary rates and gene duplication patterns among families of functional olfactory receptor genes in humans.

In humans, odors are detected by ~400 functional olfactory receptor (OR) genes. The superfamily of functional OR genes can be further divided into tens of families. In large part, the OR genes have experienced extensive tandem duplications, which have led to gene gains and losses. However, whether different OR gene families have experienced distinct modes of gene duplication has yet to be reported. We conducted comparative genomic and evolutionary analyses for human functional OR genes. Based on analysis of human-mouse 1-1 orthologs, we found that human functional OR genes show higher-than-average evolutionary rates, and there are significant differences among families of functional OR genes. Via comparison with seven vertebrate outgroups, families of human functional OR genes show different extents of gene synteny conservation. Although the superfamily of human functional OR genes is enriched in tandem and proximal duplications, there are particular families which are enriched in segmental duplications. These findings suggest that human functional OR genes may be governed by different evolutionary mechanisms and that large-scale gene duplications have contributed to the early evolution of human functional OR genes.

Humans

Large-scale analysis of MYB genes in Cucurbitaceae identifies a novel gene regulating plant height.

The MYB transcription factor (TF) family, which is involved in plant growth and development, is large and diverse. Previous studies on MYB family in Cucurbitaceae were mostly based on a single genome or focused on the R2R3 subfamily. Here, we analyzed 91 genomes of 11 Cucurbitaceae species and identified a total of 15 858 MYB genes. According to phylogenetic relationships, these genes were divided into 27 subgroups. The identified MYB genes were further classified into 121 MYB orthologous gene groups (OGGs), including 25 core, 57 softcore, 19 shell and 20 line-specific/cloud groups. Whole-genome duplication was the most common mechanism of MYB genes expansion. In core group, the higher proportions of MYB genes were found to be in the coexpression network constructed by the RNA-seq data. Through the comprehensive analysis including phylogeny and gene expression profile of cucumber MYB genes, as well as genetic variations in 103 cucumber germplasms, we identified a MYB gene CsRAX5, which may be related to cucumber plant height. We used gene editing technology to knockout and overexpress CsRAX5. In the knockout lines, Csrax5, the height was significantly increased compared with wild type (WT), whereas after overexpression the height of CsRAX5-OE plants was significantly decreased compared with WT. These results indicated that MYB gene CsRAX5 negatively regulated cucumber plant height. The large-scale analysis of MYB genes in Cucurbitaceae in this study provides insights for further investigating the evolution and function of MYB genes in Cucurbitaceae crops.

Journal Article

Integration of multi-source gene interaction networks and omics data with graph attention networks to identify novel disease genes.

MOTIVATION: The pathogenesis of diseases is closely associated with genes, and the discovery of disease genes holds significant importance for understanding disease mechanisms and designing targeted therapeutics. However, biological validation of all genes for diseases is expensive and challenging. RESULTS: In this study, we propose DGP-AMIO, a computational method based on graph attention networks, to rank all unknown genes and identify potential novel disease genes by integrating multi-omics and gene interaction networks from multiple data sources. DGP-AMIO outperforms other methods significantly on 20 disease datasets, with an average AUROC and AUPR exceeding 0.9. The superior performance of DGP-AMIO is attributed to the integration of multiomics and gene interaction networks from multiple databases, as well as triGAT, a proposed GAT-based method that enables precise identification of disease genes in directed gene networks. Enrichment analysis conducted on the top 100 genes predicted by DGP-AMIO and literature research revealed that a majority of enriched GO terms, KEGG pathways and top genes were associated with diseases supported by relevant studies. We believe that our method can serve as an effective tool for identifying disease genes and guiding subsequent experimental validation efforts. AVAILABILITY AND IMPLEMENTATION: DGP-AMIO is publicly available at https://github.com/yangkaiyuan1027/DGP-AMIO.

Gene Regulatory Networks

scPOEM: robust co-embedding of peaks and genes revealing peak-gene regulation.

MOTIVATION: Identifying regulatory elements in various chromosomal regions that influence gene expression is a fundamental challenge in epigenomics, with profound implications for understanding gene regulation and disease mechanisms. The advent of paired single-cell RNA sequencing and single-cell ATAC sequencing has created unprecedented opportunities to address this challenge by enabling simultaneous profiling of gene expression and chromatin accessibility at single-cell resolution. However, the inherent signals between them are weak due to the highly sparse and noisy nature of data. RESULTS: This article proposes single-cell meta-Path based Omics Embedding (scPOEM), a novel embedding method that jointly projects chromatin accessibility peaks and expressed genes into a shared low-dimensional space. By integrating the relationships among peak-peak, peak-gene, and gene-gene interactions, scPOEM assigns closer representations in the embedding space to related peak-gene pairs. Our experiments demonstrate that scPOEM generates stable representations of peaks and genes, outperforms existing methods in recovering biologically meaningful peak-gene regulatory relationships and enables new insights in subgroup and differential analysis of gene regulation. These results highlight its potential to uncover gene regulatory mechanisms and enhance the understanding of transcriptional regulation at single-cell resolution. AVAILABILITY AND IMPLEMENTATION: The source code of scPOEM is available at https://github.com/Houyt23/scPOEM. The datasets can be obtained from the 10&#xd7; Genomics (https://www.10xgenomics.com/datasets/pbmc-from-a-healthy-donor-granulocytes-removed-through-cell-sorting-10-k-1-standard-1-0-0) and GEO database under access codes GSE194122 and GSE239916.

Gene Expression Regulation

DyNDG: Identifying Leukemia-related Genes Based on Time-series Dynamic Network by Integrating Differential Genes.

Leukemia is a malignant disease characterized by progressive accumulation with high morbidity and mortality rates, and investigating its disease genes is crucial for understanding its etiology and pathogenesis. Network propagation methods have emerged and been widely employed in disease gene prediction, but most of them focus on static biological networks, which hinders their applicability and effectiveness in the study of progressive diseases. Moreover, there is currently a lack of special algorithms for the identification of leukemia disease genes. Here, we proposed a novel Dynamic Network-based model integrating Differentially expressed Genes (DyNDG) to identify leukemia-related genes. Initially, we constructed a time-series dynamic network to model the development trajectory of leukemia. Then, we built a background-temporal multilayer network by integrating both the dynamic network and the static background network, which was initialized with differentially expressed genes at each stage. To quantify the associations between genes and leukemia, we extended a random walk process to the background-temporal multilayer network. The results demonstrate that DyNDG achieves superior accuracy compared to several state-of-the-art methods. Moreover, after excluding housekeeping genes, DyNDG yields a set of promising candidate genes associated with leukemia progression or potential biomarkers, indicating the value of dynamic network information in identifying leukemia-related genes. The implementation of DyNDG is available at both https://ngdc.cncb.ac.cn/biocode/tool/BT7617 and https://github.com/CSUBioGroup/DyNDG.

Leukemia

Evaluating selection at intermediate scales within genes provides robust identification of genes under positive selection in M. tuberculosis clinical isolates.

Multiple studies have reported genes in the M. tuberculosis (Mtb) genome that are under diversifying selection, based on genetic variants among Mtb clinical isolates. These might reflect adaptions to selection pressures associated with modern clinical treatment of TB. Many, but not all, of these genes under selection are related to drug resistance. Most of these studies have evaluated selection at the gene-level. However, positive selection can be evaluated on different scales, including individual sites (codons) and local regions within an ORF. In this paper, we use GenomegaMap, a Bayesian method for calculating selection, to evaluate selection of genes in the Mtb genome at all three levels. We present evidence that the intermediate analysis (windows of codons) yields the most credible list of candidate genes under selection (excluding PPE and PE_PGRS genes, which are predicted less reliably due to frequent sequencing errors). A further advantage of this approach is that it identifies specific regions within proteins that are under selective pressure, which is useful for structural and functional interpretation. In an analysis of two separate collections of Mtb clinical isolates (from Moldova; and a globally-representative set), we observed 53 and 173 significant genes under selection, with 36% overlap. The lists of genes under selection include many drug-resistance genes, as well as other genes that have previously been reported to be under selection (resR, phoR). The specific regions under selection identified within drug-resistance genes are shown to correspond to protein structural features known to be involved in resistance, supporting accuracy of the method. Positive selection in several ESX-1-related genes was also observed, suggesting adaptation to immune pressure.

adaptation

Diversification of Cellulose Synthase (CESA) Genes in Mosses Suggests Both Ancient and Recent Gene duplications.

Cellulose is an important polysaccharide that constitutes all plant cell walls, giving them strength and stability. The plant cellulose synthase (CESA) gene family, which encodes the catalytic subunits of cellulose synthesis complexes (CSCs), has diversified independently in several plant lineages, providing an interesting model for understanding selection for gene duplication. Here we quantified the presence of CESA genes across mosses to understand how the process of gene family diversification occurred in this group and how it parallels diversification in other groups. We first examined the CESA gene family in eight species of mosses across seven families for which whole genome assemblies were available. We then identified CESA genes from additional species, for which only short-read sequence data was available, by using BLAST searches and targeted gene assemblies. We validated this approach by comparing the assembled paralogs from the short-read data to the genes identified from whole genome assemblies in the eight reference species. This approach allowed us to identify paralogs directly from short-read data and greatly expand our sample set. Results from the combined empirical data support the hypothesis that CESA genes diversified within the moss lineage at least as early as the mesozoic period, during or possibly even prior to the onset of moss diversification, but also continue to diversify within modern species. In addition, we found evidence for purifying selection as the dominant force shaping these genes and observed that different lineages experienced different levels of evolutionary constraint. Lastly, our approach to assemble paralogs has the potential to allow researchers to improve analyses of gene duplication events.

Physcomitrium patens

Gene behaviors-based network enrichment analysis and its application to reveal immune disease pathways enriched with COVID-19 severity-specific gene networks.

MOTIVATION: Gene network analysis is essential for understanding the complex mechanisms underlying diseases, which often involve disruptions in molecular networks rather than individual genes. Despite the availability of large-scale omics datasets and computational tools for gene network analysis, interpretation of the biological relevance of these extensive networks remains challenging. RESULTS: We propose a novel computational strategy, gene behaviors-based network enrichment analysis, which systematically identifies functional pathways enriched in phenotype-specific gene networks. Our novel method incorporates comprehensive network characteristics, i.e. gene expression levels, edge strengths, and structural patterns of edges, to rank genes based on activity and assess pathway enrichment, effectively identifying functional pathways enriched within these networks. Through simulation studies, our strategy demonstrated superior performance compared with that of existing methods in identifying enriched pathways. We applied this strategy to whole-blood RNA-seq data from 1102 COVID-19 samples provided by the Japan COVID-19 Task Force. The analysis revealed immune disease pathways enriched with COVID-19 severity-specific gene networks, including "Systemic lupus erythematosus" in asymptomatic and severe samples and "Inflammatory bowel disease," "Primary immunodeficiency," and "Rheumatoid arthritis" in mild samples. Key biomarkers of COVID-19, such as CXCL8, S100A9, and HLA class I genes, have been identified as critical hub genes and the main players within these networks. AVAILABILITY AND IMPLEMENTATION: Code is available in Figshare (https://doi.org/10.6084/m9.figshare.29093648.v3).

COVID-19

Coexpression of neighboring genes in Caenorhabditis elegans is mostly due to operons and duplicate genes.

In many eukaryotic species, gene order is not random. In humans, flies, and yeast, there is clustering of coexpressed genes that cannot be explained as a trivial consequence of tandem duplication. In the worm genome this is taken a step further with many genes being organized into operons. Here we analyze the relationship between gene location and expression in Caenorhabditis elegans and find evidence for at least three different processes resulting in local expression similarity. Not surprisingly, the strongest effect comes from genes organized in operons. However, coexpression within operons is not perfect, and is influenced by some distance-dependent regulation. Beyond operons, there is a relationship between physical distance, expression similarity, and sequence similarity, acting over several megabases. This is consistent with a model of tandem duplicate genes diverging over time in sequence and expression pattern, while moving apart owing to chromosomal rearrangements. However, at a very local level, nonduplicate genes on opposite strands (hence not in operons) show similar expression patterns. This suggests that such genes may share regulatory elements or be regulated at the level of chromatin structure. The central importance of tandem duplicate genes in these patterns renders the worm genome different from both yeast and human.

Animals

Differential gene expression study in whole blood identifies candidate genes for psychosis in African American individuals.

Genome-wide association has identified regions of the genome that mediate risk for psychosis. It is possible that variants in these regions confer risk by altering gene expression. This work has predominantly been conducted in individuals of European descent and has focused narrowly on schizophrenia rather than psychosis as a syndrome. In the present study we investigated alterations in gene expression in African American individuals with a range of psychotic diagnoses to increase understanding of the etiology in an underserved population. We performed RNA-seq in whole bloody to survey the transcriptome in 126 patients with a psychosis-spectrum disorder and 217 healthy controls and applied differential gene expression analyses across the genome while controlling for age, sex, population stratification and batch. We found 18 differentially expressed genes (DEGs), some of the locations of the corresponding genes overlap with previously implicated regions for psychosis, but many of which were novel associations. Enrichment analysis of nominally significant genes (p&#xa0;<&#xa0;0.05) revealed overrepresentation of biological processes relating to platelet, immune and cellular function, and sensory perception. Weighted gene co-expression network analysis, applied to identify modules of co-expressed genes associated with psychosis, revealed 10 modules, one of which was significantly associated with psychosis. This module was significantly enriched for DEGs, and for platelet function. These results support the potential role of immune function in the etiology of psychosis, identify novel candidate gene expression phenotypes that correspond to both established and new genomic regions, in individuals of African American ancestry.

Humans

Recent gene duplication and structural remodeling drive rapid lineage-specific gene family evolution in plants.

Gene duplication promotes the generation of novel gene functions and trait diversity across species. Here, we present DupHIST, a computational pipeline that reconstructs the hierarchical timing of gene duplications by integrating maximum likelihood (ML)-based phylogeny with substitution-derived timing via statistical smoothing. Applied to over 4.5 million genes from 114 plant genomes, we successfully inferred duplication histories across nearly 130,000 orthogroups. This large-scale analysis showed that 53.0% of genes arose from recent, lineage-specific duplications, with high concentrations in particular multi-copy families. Among these, NLR, C48, and P450 families exemplified how recently duplicated genes undergo rapid stepwise structural remodeling. This process was primarily driven by small-scale mutations, including insertions, deletions, and frameshifts, that rapidly accumulated shortly after duplication. By resolving the precise duplication order, we reconstructed these architectural changes, thereby enabling both the inference of putative ancestral structures and the exploration of functional diversification arising from structural remodeling. Structure-based clustering further uncovered that recently duplicated, uncharacterized genes retain core domain structures resembling known functional proteins even across phylogenetically distant species lacking sequence homology. Our findings reveal that recent gene duplications and subsequent structural remodeling represent a widespread and lineage-specific force driving rapid diversification of gene families in plants.

Gene duplication history

The msf gene causes condition-specific shifts in global gene expression in Haemophilus influenzae.

UNLABELLED: Haemophilus influenzae is a diverse human-restricted bacterium that normally colonizes the healthy nasopharynx but also causes common infections. Comparisons of clinical isolate genomes previously identified a gene, msf, that contained Sel1-like repeats that were associated with clinical disease. Mutant analysis had further found that msf improved survival in macrophages and increased systemic infection in an animal model. However, the role of msf in other conditions and its molecular function remain unknown. To identify protein-protein interactions with Msf, a yeast two-hybrid screen against an H. influenzae prey library was conducted, which found potential interactions with lipoprotein exporter protein LolD and an autotransporter adhesin Hap. To identify effects of msf on gene expression, we compared wild-type and mutant strains grown in multiple culture conditions by RNA-seq. The results indicate that msf modulates global gene expression in a condition-dependent manner, exerting an especially strong influence in starved surface-attached biofilm cells. The few consistent changes in mutants' planktonic exponential and stationary phases included decreased expression of two paralogous autotransporter adhesins. By contrast, mutant cells in starved surface-attached biofilms had dramatic changes in expression, including upregulation of protein translation and downregulation of alternative carbon metabolism. However, assays of 24 hour biofilm phenotypes found only subtle gene expression changes. Together, the results point to a speculative model of Msf functioning as an envelope-associated chaperone whose presence affects the relative expression of proteins at the outer membrane. IMPORTANCE: Comparing genomes from different clinical isolates of the same pathogenic bacterial species has identified genes associated with virulence, but many of these are understudied or have no known function. The msf gene was previously implicated as a virulence factor in Haemophilus influenzae, a common cause of mucosal diseases including middle-ear and chronic lung infections. This study finds that the msf gene causes condition-specific changes in gene expression, with especially dramatic changes in starved surface-attached biofilm cells. Along with identification of putative protein-protein interaction partners, the results provide new clues as to the molecular and cellular function of Msf, potentially as an envelope-associated chaperone involved in membrane protein trafficking. Understanding how virulence-associated genes like msf modulate bacterial responses to the environment may help explain why some bacterial strains remain harmless colonizers while others become pathogens.

Haemophilus influenzae

Gene regulation technologies for gene and cell therapy.

Gene therapy stands at the forefront of medical innovation, offering unique potential to treat the underlying causes of genetic disorders and broadly enable regenerative medicine. However, unregulated production of therapeutic genes can lead to decreased clinical utility due to various complications. Thus, many technologies for controlled gene expression are under development, including regulated transgenes, modulation of endogenous genes to leverage native biological regulation, mapping and repurposing of transcriptional regulatory networks, and engineered systems that dynamically react to cell state changes. Transformative therapies enabled by advances in tissue-specific promoters, inducible systems, and targeted delivery have already entered clinical testing and demonstrated significantly improved specificity and efficacy. This review highlights next-generation technologies under development to expand the reach of gene therapies by enabling precise modulation of gene expression. These technologies, including epigenome editing, antisense oligonucleotides, RNA editing, transcription factor-mediated reprogramming, and synthetic genetic circuits, have the potential to provide powerful control over cellular functions. Despite these remarkable achievements, challenges remain in optimizing delivery, minimizing off-target effects, and addressing regulatory hurdles. However, the ongoing integration of biological insights with engineering innovations promises to expand the potential for gene therapy, offering hope for treating not only rare genetic disorders but also complex multifactorial diseases.

Humans

A duplicated female pathway gene figla-like evolves as the male sex-determining gene in tilapia.

As the largest group of vertebrates, fish exhibit frequent turnover of sex-determining (SD) genes. Here, we assemble a chromosome-level YY red tilapia genome and identify figla-like (figlal) as the SD gene on tilapia linkage group (LG) 1. Integrative phylogenetic and genomic evidence suggests that figlal originated from a tilapia-specific duplication and transposition of the ancestral bHLH family gene figla from LG12 to LG1. Fluorescence in situ hybridization reveals expression divergence between figla and figlal, with figla expressed in female oocytes and figlal expressed in male gonadal somatic cells during early gonadal differentiation. The shift in expression after duplication might be driven by the insertion of cis-regulatory elements mediated by transposable elements. Knockout of figlal in XY fish results in male-to-female sex reversal as indicated by ovarian morphology, down-regulation of the male pathway gene dmrt1, and up-regulation of the female pathway gene cyp19a1a in the gonads. In contrast, overexpression of figlal in XX fish induces female-to-male sex reversal. These findings implicate figlal as an SD gene on tilapia LG1 and reveal the history of a unique evolutionary innovation in which a female oocyte gene evolved into a male SD gene via duplication, transposition, and cis-regulatory rewiring.

Animals