PubMed HealthSearch

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Unravelling sex differences in the genetic architecture of anxiety.

BACKGROUND: Anxiety disorders show striking sex differences in prevalence, symptoms, and clinical characteristics, shaping how they manifest and are experienced. METHODS: Here, we report the first sex-specific meta-analysis of genome-wide association studies (GWAS) of anxiety, leveraging two of the largest biobank datasets, UK Biobank and All of Us, comprising 85,042 female cases with 196,789 controls and 36,732 male cases with 136,924 controls. Functional annotation, sex-specific polygenic scores (PGS), and genetic correlations were performed to assess genetic differences and functional implications. RESULTS: In females, 21 lead SNPs were significantly associated with anxiety, compared to five in males. Although the genetic correlation between sexes was high, it was significantly different from one, indicating partially distinct genetic architectures. In addition, both the SNP-based observed and liability-scale heritabilities (assuming a 2:1 female-to-male prevalence ratio) were significantly higher in females. Gene-based tests and functional prioritization identified different genes associated with anxiety in females and males. Moreover, genetic correlation analyses revealed stronger associations of female anxiety with attention-deficit/hyperactivity disorder (ADHD) and body mass index (BMI), whereas male anxiety showed stronger correlations with waist-hip-ratio-adjusted BMI. CONCLUSIONS: While the overall genetic architecture of anxiety is largely shared, our findings reveal distinct sex-specific genetic associations and correlations, highlighting the value of analyzing the sexes separately to uncover genetic signals that may be masked in sex-combined samples.

Female

Uniform processing and analysis of IGVF massively parallel reporter assay data with MPRAsnakeflow.

As researchers and clinicians seek to identify human genomic alterations relevant to traits and disorders, identifying and aggregating evidence providing mechanistic support for associations between alterations and phenotypes remains challenging. In particular, the study of noncoding genomic variation remains a major challenge because of the lack of accurate functional annotation for activity in a given context and across alleles. Experimental evidence is critical for prioritizing and interpreting functional effects of genetic alterations. Massively parallel reporter assays (MPRAs) have emerged as a powerful high-throughput approach, enabling quantification of regulatory element activity and allelic effects, as well as systematic dissection of gene regulatory logic and variant effects across different contexts. However, the diversity of MPRA designs, lack of standardized formats, and many potential processing parameters hamper data integration, reproducibility, and meta-analyses across studies. To address these challenges, the Impact of Genomic Variation on Function (IGVF) Consortium established an MPRA focus group to develop community standards, including harmonized file formats, and robust analysis pipelines for a wide range of library types and experimental designs. Here, we present these formats and comprehensive computational tools, MPRAlib and MPRAsnakeflow, for uniform processing from raw sequencing reads to counts, processing, and visualization. Using diverse MPRA data sets, we investigated technical variability sources including barcode sequence bias, outlier barcodes, and delivery method (episomal vs. lentiviral). Our results establish best practices for MPRA data generation and analysis, facilitating robust, reproducible research and large-scale integration. The presented tools and standards are publicly available, providing a foundation for future collaborative efforts in regulatory genomics.

Humans

A chromosome-level reference genome assembly of the Small snakehead (Channa asiatica).

The Small snakehead (Channa asiatica) is an economically important species in both aquaculture and ornamental trade, mainly distributed in South China and Southeast Asia. Despite its significance, limited genomic resources have impeded in-depth genetic studies and breeding programs. In this study, we used PacBio HiFi long-read sequencing, Illumina short-read sequencing, and Hi-C technologies to generate a high-quality chromosome-level genome of the C. asiatica. The final genome spans 659.44 Mb, with an impressive 98.18% anchored to 23 chromosomes. Notably, the contig N50 and scaffold N50 are 23.92 Mb and 29.61 Mb, validated by a BUSCO completeness score of 98.93%. Genome annotation identified 26,603 protein-coding genes, 99.29% of which were confirmed by BUSCO analysis, and 93.68% were functionally annotated. Approximately 27.72% of the genome sequences were classified as repeat elements. This high-fidelity genome assembly provides a robust foundation for advancing molecular breeding, comparative genomics, and evolutionary studies of C. asiatica and related species.

Animals

Systematic discovery of retina-enriched Rik genes identifies 1190005I06Rik as a novel modulator of visual signalling.

BACKGROUND: High‑throughput transcriptome projects have revealed thousands of mammalian genes with little or no functional annotation. Among these are hundreds of loci assigned provisional “Rik” identifiers following discovery in the RIKEN cDNA annotation effort. Although often dismissed as genomic dark matter, such genes may encode tissue‑restricted proteins that modulate physiologic functions and influence disease. The retina is a highly specialised neural tissue and a common site of inherited disorders; understanding its molecular repertoire could illuminate novel therapeutic avenues. METHODS: We integrated bulk RNA‑seq from ten adult mouse tissues, evolutionary and domain analysis, single‑cell RNA‑seq, and CRISPR/Cas9 gene disruption to systematically catalogue protein‑coding Rik genes enriched in the retina and test the function of a representative gene. RESULTS: A rigorous differential expression analysis identified 44 Rik genes with robust retina‑specific expression compared with nine non‑retinal tissues. Many of these genes lack orthologues beyond rodents, while others show broad conservation, illustrating a continuum from lineage‑restricted to conserved retinopathy candidates. Single‑cell transcriptomics revealed that these genes are expressed across retinal cell types, with the highest aggregate expression in cone photoreceptors and inner interneurons. To evaluate physiological significance, we generated a 1190005I06Rik knockout mouse. Although retinal architecture appeared normal, loss of 1190005I06Rik enhanced electroretinogram b‑wave amplitudes and altered light‑avoidance behaviour, indicating that this previously uncharacterised gene acts as a negative modulator of visual signalling. CONCLUSIONS: We present a curated atlas of retina‑enriched Rik genes and demonstrate that 1190005I06RIK modulates retinal circuit function. This resource expands the molecular landscape of the retina and provides new candidates for the genetic basis of inherited retinal disease. Our findings underscore that unannotated genes may exert measurable effects on sensory processing and warrant systematic exploration in the context of human ocular disorders.

Animals

DeepES: deep learning-based enzyme screening to identify orphan enzyme genes.

MOTIVATION: Progress in sequencing technology has led to determination of large numbers of protein sequences, and large enzyme databases are now available. Although many computational tools for enzyme annotation were developed, sequence information is unavailable for many enzymes, known as orphan enzymes. These orphan enzymes hinder sequence similarity-based functional annotation, leading gaps in understanding the association between sequences and enzymatic reactions. RESULTS: Therefore, we developed DeepES, a deep learning-based tool for enzyme screening to identify orphan enzyme genes, focusing on biosynthetic gene clusters and reaction class. DeepES uses protein sequences as inputs and evaluates whether the input genes contain biosynthetic gene clusters of interest by integrating the outputs of the binary classifier for each reaction class. The validation results suggested that DeepES can capture functional similarity between protein sequences, and it can be implemented to explore orphan enzyme genes. By applying DeepES to 4744 metagenome-assembled genomes, we identified candidate genes for 236 orphan enzymes, including those involved in short-chain fatty acid production as a characteristic pathway in human gut bacteria. AVAILABILITY AND IMPLEMENTATION: DeepES is available at https://github.com/yamada-lab/DeepES. Model weights and the candidate genes are available at Zenodo (https://doi.org/10.5281/zenodo.11123900).

Deep Learning

The current and future perspective of ChickenGTEx project and its applications in precision breeding.

The Chicken Genotype-Tissue Expression (ChickenGTEx) project was established to systematically characterize the regulatory landscape of the chicken genome and to accelerate the translation of functional genomics into precision breeding. By integrating whole-genome sequencing with multi-tissue transcriptomic profiling, ChickenGTEx provides a comprehensive atlas of gene expression regulation across diverse tissues and physiological systems. Current findings demonstrate that complex production traits are governed by coordinated regulatory networks rather than isolated loci, with substantial contributions from tissue-specific gene expression, structural variation, and genotype-by-sex interactions. Sex-dependent regulatory effects further refine the genetic architecture of metabolic, immune, and reproductive traits, highlighting the importance of incorporating sex as a biological variable in genomic analyses. Application of integrative omics frameworks within elite layer populations has revealed multilayer regulatory mechanisms underlying extended laying performance, feed efficiency, metabolic health, and eggshell quality. By partitioning phenotypic variance into genetic, regulatory, and host-microbiome components, these approaches move beyond association-based mapping toward causal inference and biological interpretation. Importantly, validated regulatory loci identified through ChickenGTEx and related analyses provide actionable markers for genomic selection and rational targets for precision genome modification. Looking forward, continued expansion of regulatory atlases, incorporation of single-cell and longitudinal data in diverse environmental conditions, and integration of functional annotation into breeding pipelines will further enhance prediction accuracy and sustainable genetic improvement. The ChickenGTEx project thus represents a foundational platform bridging functional genomics and practical poultry breeding.

Animals

Chromosome-level genome assembly of the hemiparasitic Taxillus sutchuenensis (Loranthaceae).

Taxillus sutchuenensis, an ecologically and medicinally important hemiparasitic plant that parasitizes diverse woody hosts, was sequenced to generate a high-quality chromosome-level genome assembly. PacBio HiFi long reads, RNA-seq transcriptome data, and Hi-C data were used to assemble a 406.32 Mb genome anchored onto nine pseudo-chromosomes, with a scaffold N50 of 45.59 Mb. The assembly showed high completeness and accuracy, supported by BUSCO (93.6%) and Merqury QV (70.6) assessments. The LTR Assembly Index (LAI) of 13.98 indicated excellent continuity. A total of 21,795 protein-coding genes were predicted, with 94.46% functionally annotated. Repetitive sequences accounted for 50.05% of the genome, primarily LTR retrotransposons. This genome provides a valuable resource for investigating the evolution, functional genomics, and parasitic mechanisms of hemiparasitic plants.

Genome, Plant

Accounting for recombination rate variation improves inference of barrier loci and reveals the role of both natural and sexual selection in an incipient bird radiation.

Examining genomic patterns of differentiation across lineage pairs at different stages of the speciation continuum, in combination with recombination maps, can help disentangle the effects of linked and divergent selection and identify lineage-specific targets of selection that may act as barrier loci during speciation. Here, we apply this framework to genomic data from African and Indian Ocean bird species of the genus Zosterops (Zosteropidae) to identify candidate barrier loci between ecologically, phenotypically, and genetically distinct Reunion gray white-eye (Zosterops borbonicus) parapatric geographic forms. Using analyses that account for recombination rate variation, we show that putative targets of divergent selection are primarily located on the Z chromosome, except in comparisons between geographic forms that differ in their ecologies. Functional annotation revealed that candidate barrier loci between forms with similar environmental niches are associated with genes involved in song formation and immune function, whereas those between forms with different environmental niches are associated with adaptation to altitude, morphology, and song behavior. Our results highlight the combined roles of natural and sexual selection in the evolution of reproductive barriers in this incipient species radiation.

Animals

Comparative genomics reveals population structure and functional differentiation in Limosilactobacillus fermentum.

Limosilactobacillus fermentum is a widely distributed lactic acid bacterium frequently detected in fermented foods and host-associated microbiota, yet its global genomic diversity and functional variability remain insufficiently characterized. Here, we performed a large-scale comparative genomic analysis of 336 high-quality L. fermentum genomes curated from public databases. Species identity was validated using average nucleotide identity (ANI), and population structure was examined using pairwise ANI comparisons together with Mash-based phylogenetic reconstruction. Clustering at ≥ 99% ANI resolved the dataset into 15 genomic clusters, with four dominant lineages comprising the majority of genomes. Pangenome reconstruction identified 5,853 gene clusters, including 1,325 core genes (22.6%) and a large accessory component dominated by low-frequency genes. Heap's law modeling (λ = 0.19) indicated a weakly open pangenome, suggesting ongoing gene acquisition as additional genomes are sampled. Functional annotation revealed that core genes were primarily associated with essential cellular processes, whereas accessory genes were enriched in carbohydrate metabolism, membrane-associated functions, and defense-related systems. Variation in carbohydrate-active enzymes (CAZymes), transport systems, and stress-response genes was observed across lineages, indicating strain-level functional diversity. Although genomes from human and food sources were broadly distributed across phylogenetic lineages, multivariate analysis showed that gene-content variation was more strongly associated with genomic lineage than with isolation source. These results provide a population genomic framework for understanding genomic diversity and functional potential in L. fermentum.

Phylogeny

Chromosome-level genome assembly of Sinocyclocheilus jii based on PacBio HiFi and Hi-C sequencing.

Sinocyclocheilus jii, a cavefish species endemic to China, belongs to the genus Sinocyclocheilus within the family Cyprinidae. Species within this genus exhibit significant morphological differentiation, making it not only the most species-rich genus within Cyprinidae in China but also the most diverse group of cavefishes worldwide. However, the limited availability of genomic resources has limited investigations into the genetic basis of trait variations, phylogenetic relationships, and adaptive evolution in this genus. In this study, we assembled a chromosome-level reference genome for S. jii by integrating PacBio HiFi long reads, Illumina short reads, and Hi-C sequencing data. Flow cytometry was used to estimate the genome size prior to assembly, providing a key step in technical validation. The final genome assembly spans 1.75 Gb with a contig N50 of 35.0 Mb. Using Hi-C sequencing data, the assembled scaffolds were successfully anchored to 50 chromosomes. The completeness of the chromosome-level assembly was estimated at 98.9% by BUSCO analysis. Genome annotation identified 855.5 Mb of repetitive sequences and predicted a total of 52,867 protein-coding genes, of which 51,932 genes were functionally annotated. This study presents a high-quality chromosome-level genome assembly and annotation of S. jii, providing a fundamental genomic resource for future phylogenetic and evolutionary studies.

Animals

De novo assembly of transcriptomes of six Hua species (Semisulcospiridae, Cerithioidea, Gastropoda).

Species in Semisulcospiridae are important in freshwater ecology and have great research value, yet their genomic resources remain very limited. Here, we present de novo assembled transcriptomes from six species of Hua in Semisulcospiridae, including Hua textrix (Heude, 1888), H. yangi L.-N. Du, J.-X. Yang & Chen, 2023, H. wujiangensis L.-N. Du, J.-X. Yang & Chen, 2023, and three undescribed species. Assembly was performed using Trinity, resulting in average contig lengths ranging from 716.6 to 883.3 bp and transcript numbers ranging from 147,147 to 268,741. Benchmarking Universal Single-Copy Ortholog (BUSCO) analysis was used to assess the transcriptome completeness. The functional annotation of transcripts for each species had over 18,000 BLAST hits, 17,000 GO terms, 15,000 KEGG pathways, 8,000 Pfam accessions, and 140 COG functional categories. This study provides valuable transcriptomic resources for the six Hua species, which can be used for various research of Semisulcospiridae, including biodiversity, phylogeny, and comparative genomics.

Transcriptome

Functional screening of ZIP8 naturally occurring variants identifies pathogenic mutations and trafficking defects.

The rapid expansion of human genomic data has revealed a large number of naturally occurring variants, creating a major challenge for functional annotation. The human metal transporter SLC39A8 (ZIP8) is a clinically important divalent metal transporter, yet most of its documented variants remain uncharacterized. Here, we developed a workflow to functionally evaluate ZIP8 variants by integrating laser ablation inductively coupled plasma time-of-flight mass spectrometry (LA-ICP-TOF-MS) with scaled-up cell-based transport assays. Using this method, we systematically analyzed 33 naturally occurring missense variants located in the extracellular domain (ECD) of ZIP8. The assay enables direct quantification of intracellular metal accumulation with substantially improved throughput (∼150 samples per hour). Functional screening identified 14 potential pathogenic variants with significantly reduced transport activity. Comparison with computational predictions revealed a moderate correlation between activity and AlphaMissense pathogenicity scores (R2 = 0.423), while an error rate of ∼20% for AlphaMissense underscores the need for experimental validation. Flow cytometry analysis showed that most loss-of-function variants exhibit impaired trafficking of the protein to the cell surface possibly due to mutation-caused protein misfolding or instability. Structural mapping of activity-compromised variants, together with functional assessment of the ZIP8-ECD, highlights the importance of this domain in ZIP8 expression and intracellular protein trafficking. Together, this work establishes a scalable approach for functional screening of metal transporter variants and provides new insights into the structure-function relationships of ZIP8.

Journal Article

Dental wastewater reveals a hidden reservoir of oral bacteriophage diversity.

Bacteriophages (phages) are being explored as alternatives or complements to antibiotics because of their ability to selectively kill bacterial pathogens. However, phages that infect many oral bacteria remain undiscovered. Here, we discovered that dental wastewater harbors previously underexplored phage diversity. Viral particles concentrated from dental wastewater displayed diverse morphologies, including abundant filamentous phage-like particles. Deep long-read metagenomic sequencing of concentrated viral particles generated 7.4 billion bases of sequence data and yielded 255 medium- to high-quality viral operational taxonomic units (vOTUs), including 46 predicted complete genomes. Comparison with large phage databases revealed that 63 of these 255 vOTUs had no detectable match, indicating that extensive sequencing of dental wastewater substantially expands the number of potential bacteriophages associated with the human oral microbiome. Host prediction linked many vOTUs to oral-associated bacterial taxa, including species with few or no previously reported phages, such as Porphyromonas gingivalis, Tannerella forsythia, and Candidatus Saccharibacteria. Functional annotation identified diverse genes associated with antiphage defense systems within a subset of vOTUs, suggesting that oral phages may contribute to the movement of genes encoding bacterial immune functions within the oral microbiome. Together, these findings expand the known oral phageome and show that dental wastewater contains a largely untapped diversity of phages.IMPORTANCEThe human oral cavity contains a diverse microbial community, but the bacteriophages (phages) that infect many oral bacteria remain poorly characterized. This gap limits our understanding of how phages shape oral microbial communities. Here, we show that dental wastewater is an underexplored source of oral phage diversity. Deep long-read metagenomic sequencing revealed 255 medium- to high-quality phage operational taxonomic units, many of which are not present in existing oral phage databases. These genomes include predicted phages of periodontal disease-associated bacteria and other oral taxa with few or no known phages. Dental wastewater therefore expands the known human oral phageome and reveals candidate phages linked to bacteria associated with oral health and disease.

Bacteriophages

AnoEST: toward A. gambiae functional genomics.

Here, we present an analysis of 215,634 EST and cDNA sequences of a major vector of human malaria Anopheles gambiae structured into the AnoEST database. The expressed sequences are grouped into clusters using genomic sequence as template and associated with inferred functional annotation, including the following: corresponding Ensembl gene prediction, putative orthologous genes in other species, homology to known proteins, protein domains, associated Gene Ontology terms, and corresponding classification into broad GO-slim functional groups. AnoEST is a vital resource for interpretation of expression profiles derived using recently developed A. gambiae cDNA microarrays. Using these cDNA microarrays, we have experimentally confirmed the expression of 7961 clusters during mosquito development. Of these, 3100 are not associated with currently predicted genes. Moreover, we found that clusters with confirmed expression are nonbiased with respect to the current gene annotation or homology to known proteins. Consequently, we expect that many as yet unconfirmed clusters are likely to be actual A. gambiae genes. [AnoEST is publicly available at http://komar.embl.de, and is also accessible as a Distributed Annotation Service (DAS).].

Animals

The Mycobacterium tuberculosis Transposon Sequencing Database (MtbTnDB): A Large-Scale Guide to Genetic Conditional Essentiality.

Characterizing genetic essentiality across various conditions is fundamental for understanding gene function. Transposon sequencing (TnSeq) is a powerful technique to generate genome-wide essentiality profiles in bacteria and has been extensively applied to Mycobacterium tuberculosis (Mtb). Dozens of TnSeq screens have yielded valuable insights into the biology of Mtb in vitro, inside macrophages, and in model host organisms. Despite their value, these Mtb TnSeq profiles have not been standardized or collated into a single, easily searchable database. This results in significant challenges when attempting to query and compare these resources, limiting our ability to obtain a comprehensive and consistent understanding of genetic conditional essentiality in Mtb. We address this problem by building a central repository of publicly available Mtb TnSeq screens, the Mtb transposon sequencing database (MtbTnDB). The MtbTnDB is a living resource that encompasses to date ≈150 standardized TnSeq screens, enabling open access to data, visualizations, and functional predictions through an interactive web app (www.mtbtndb.app). We conduct several statistical analyses on the complete database, such as demonstrating that (i) genes in the same genomic neighborhood have similar TnSeq profiles, and (ii) clusters of genes with similar TnSeq profiles are enriched for genes from similar functional categories. We further analyze the performance of machine learning models trained on TnSeq profiles to predict the functional annotation of orphan genes in Mtb. By facilitating the comparison of TnSeq screens across conditions, the MtbTnDB will accelerate the exploration of conditional genetic essentiality, provide insights into the functional organization of Mtb genes, and help predict gene function in this important human pathogen.

DNA Transposable Elements

Functional screening of ZIP8 naturally occurring variants identifies pathogenic mutations and trafficking defects.

The rapid expansion of human genomic data has revealed a large number of naturally occurring variants, creating a major challenge for functional annotation. The human metal transporter SLC39A8 (ZIP8) is a clinically important, promiscuous divalent metal transporter, yet most of its documented variants remain uncharacterized. Here, we developed a workflow to functionally evaluate ZIP8 variants by integrating laser ablation inductively coupled plasma time-of-flight mass spectrometry (LA-ICP-TOF-MS) with scaled-up cell-based transport assays. Using this method, we systematically analyzed 33 naturally occurring missense variants located in the extracellular domain (ECD) of ZIP8. The assay enables direct quantification of intracellular metal accumulation with substantially improved throughput (~150 samples per hour). Functional screening identified 14 potential pathogenic variants with significantly reduced transport activity. Comparison with computational predictions revealed a moderate correlation between activity and AlphaMissense pathogenicity scores (R2 = 0.423), while an error rate of ~20% underscores the need for experimental validation. Flow cytometry analysis showed that most loss-of-function variants exhibit impaired trafficking of the protein to the cell surface possibly due to mutation-caused protein misfolding or instability. Structural mapping of activity-compromised variants, together with functional assessment of the ZIP8-ECD, highlights the importance of this domain in ZIP8 expression and intracellular trafficking. Together, this work establishes a scalable approach for functional screening of metal transporter variants and provides new insights into the structure-function relationships of ZIP8.

Journal Article

Genome-Wide Identification and Characterization of the TBL Gene Family and Temporal Expression Dynamics During Powdery Mildew Infection in Cucumber (Cucumis sativus).

Cell-wall polysaccharide O-acetylation contributes to cell-wall assembly, organ development, and plant-pathogen interactions, but the cucumber TBL gene family remains poorly characterized. Here, 37 CsTBL genes were identified genome-wide and analyzed using phylogenetic, syntenic, conserved-motif, gene-structure, promoter, protein-structure, Gene Ontology, and transcriptome approaches, followed by RT-qPCR analysis after powdery mildew inoculation. All CsTBL proteins contained the conserved GDS and DxxH motifs, whereas accessory motifs and predicted structural features varied among clades. Intraspecific analysis identified dispersed, WGD/segmental, and tandem duplication categories, and cross-species synteny was more extensive with melon than with Arabidopsis. Homology-derived annotations associated CsTBL genes with cell-wall polysaccharide metabolism, Golgi/endomembrane compartments, and O-acetyltransferase activity, including six genes assigned to xylan O-acetyltransferase-related annotations. Expression profiling revealed tissue- and developmental-stage-dependent patterns, whereas the publicly available powdery mildew RNA-seq dataset provided descriptive temporal expression profiles in Podosphaera xanthii-inoculated samples. Independent RT-qPCR analysis using time-matched mock controls revealed distinct post-inoculation responses among six selected genes. Relative to the corresponding mock controls, CsTBL2 was consistently repressed; CsTBL15 showed transient induction at 1 dpi followed by repression; CsTBL24 exhibited a biphasic response; CsTBL25 was induced at all sampled post-inoculation time points; CsTBL26 showed progressive induction; and CsTBL30 reached its highest observed expression level at 3 dpi. Integrated functional annotation and expression evidence highlighted CsTBL26 as a priority candidate for further functional characterization, while CsTBL24 and CsTBL25 represented fruit-associated candidates with distinct powdery mildew responses; CsTBL30 remained an additional strongly infection-responsive candidate. These findings provide an evolutionary and expression-based framework for the functional characterization of the cucumber TBL gene family.

O-acetylation

Genetic overlap between depression and C-reactive protein levels: Evidence from a cross-trait analysis.

Inflammation and depression have been consistently associated, with elevated C-reactive protein (CRP) levels observed in a significant subset of affected individuals. However, the genetic mechanisms underlying this association remain poorly understood. We integrated results from large-scale genome-wide association studies (GWAS) of depression and CRP levels in a cross-trait analysis specifically focusing on identifying horizontally pleiotropic loci. Identified variants were stratified as concordant versus discordant based on their direction of effects on the two traits and followed up using functional annotation, gene set enrichment, and colocalization analyses. We also explored causal relationships using Mendelian Randomization (MR) analysis with extensive sensitivity analyses, including adjustment for body mass index (BMI). We identified 9 novel loci. Functional analyses revealed that concordant loci were enriched in genes linked to immune and inflammatory processes, while discordant loci mostly mapped to metabolic pathways, including lipid regulation. MR provided strong evidence for body mass index driving a causal relationship between the genetic liability of depression on CRP levels. Our findings suggest that the association between depression and CRP levels is partly driven by shared genetic influences, pointing to different biological pathways depending on whether genetic effects are concordant or discordant. These results underscore the importance of considering effect direction when assessing the genetic overlap between depression and inflammatory processes. In addition, they highlight BMI as a key factor in the causal relationship between depression and systemic inflammation.

C-Reactive Protein