PubMed HealthSearch

SEARCH · PubMed Health

Results for “GWAS”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Multidimensional GWAS analyses on longitudinal phenotypes reveal candidate genes regulating multi-stage egg production traits in Wannan yellow chicken.

Egg production performance directly determines the economic viability of indigenous chicken breeding. However, the genetic regulation of multi-stage egg production traits remains difficult to characterize due to their complex and dynamic nature. Here, we integrated a multidimensional GWAS framework, including single-trait GWAS, multi-trait GWAS (MTAG), and longitudinal trajectory-based GWAS (TrajGWAS), to identify stage-specific and shared genetic effects underlying egg production traits in Wannan yellow chickens (WNY). Whole-genome sequencing of 354 WNY hens (10× depth) and quality control yielded 14,253,816 SNPs for analysis. Selective sweep analyses comparing red jungle fowl, commercial layers, and WNY identified a genomic region containing IGF1 under significant selection pressure. Single-trait GWAS identified SNPs 4_57990480 (BMPR1B) and 17_370912 (LOC112531479) associated with egg production across three laying stages (21-30, 31-40, and 21-40 weeks). MTAG further identified loci 8_4336468 (FASLG) and 21_654726 (CHD5) with shared effects across the laying period, whereas TrajGWAS revealed longitudinal associations involving PRKG1 and identified dynamic loci associated with clutch traits, including GRID1. For clutch traits, stage-specific loci were detected for average clutch size (ACS) and maximum clutch size (MCS), including SNP 8_8542036 at 21-30 weeks, PROK1 at 31-40 weeks, and CUL5, ALKBH8 across the entire laying period. These results demonstrate that integrating complementary GWAS strategies improves the resolution of genetic architecture underlying egg production traits by capturing trait-specific, shared, and stage-dependent genetic effects. The identified GWAS loci and selective-sweep candidate regions provide insights into the genetic architecture of egg production traits and breed differentiation.

Egg production

Gene co-expression analysis identifies brain regions and cell types involved in migraine pathophysiology: a GWAS-based study using the Allen Human Brain Atlas.

Migraine is a common disabling neurovascular brain disorder typically characterised by attacks of severe headache and associated with autonomic and neurological symptoms. Migraine is caused by an interplay of genetic and environmental factors. Genome-wide association studies (GWAS) have identified over a dozen genetic loci associated with migraine. Here, we integrated migraine GWAS data with high-resolution spatial gene expression data of normal adult brains from the Allen Human Brain Atlas to identify specific brain regions and molecular pathways that are possibly involved in migraine pathophysiology. To this end, we used two complementary methods. In GWAS data from 23,285 migraine cases and 95,425 controls, we first studied modules of co-expressed genes that were calculated based on human brain expression data for enrichment of genes that showed association with migraine. Enrichment of a migraine GWAS signal was found for five modules that suggest involvement in migraine pathophysiology of: (i) neurotransmission, protein catabolism and mitochondria in the cortex; (ii) transcription regulation in the cortex and cerebellum; and (iii) oligodendrocytes and mitochondria in subcortical areas. Second, we used the high-confidence genes from the migraine GWAS as a basis to construct local migraine-related co-expression gene networks. Signatures of all brain regions and pathways that were prominent in the first method also surfaced in the second method, thus providing support that these brain regions and pathways are indeed involved in migraine pathophysiology.

Atlases as Topic

BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets.

MOTIVATION: statistics from genome-wide association studies (GWAS) are widely used in fine-mapping and colocalization analyses to identify causal variants and their enrichment in functional contexts, such as affected cell types and genomic features. With the expansion of functional genomic (FG) datasets, which now include hundreds of thousands of tracks across various cell and tissue types, it is critical to establish scalable algorithms integrating thousands of diverse FG annotations with GWAS results. RESULTS: We propose BTS (Bayesian Tissue Score), a novel, highly efficient algorithm uniquely designed for (i) identifying affected cell types and functional elements (context-mapping) and (ii) fine-mapping potentially causal variants in a context-specific manner using large collections of cell type-specific FG annotation tracks. BTS leverages GWAS summary statistics and annotation-specific Bayesian models to analyze genome-wide annotation tracks, including enhancers, open chromatin, and histone marks. We evaluated BTS on GWAS summary statistics for immune and cardiovascular traits, such as Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis (RA), Systemic Lupus Erythematosus (SLE), and Coronary Artery Disease (CAD). Our results demonstrate that BTS is over 100× more efficient in estimating functional annotation effects and context-specific variant fine-mapping compared to existing methods. Importantly, this large-scale Bayesian approach prioritizes both known and novel annotations, cell types, genomic regions, and variants and provides valuable biological insights into the functional contexts of these diseases. AVAILABILITY AND IMPLEMENTATION: Docker image is available at https://hub.docker.com/r/wanglab/bts with preinstalled BTS R package (https://bitbucket.org/wanglab-upenn/BTS-R) and BTS GWAS summary statistics analysis pipeline (https://bitbucket.org/wanglab-upenn/bts-pipeline).

Genome-Wide Association Study

A module-based approach for post-omics, post-GWAS network-based gene classification.

MOTIVATION: Complex traits and diseases are highly polygenic and understanding the full set of genes involved is a central challenge in biomedicine. However, due to sample size limitations and noise (technical and biological), experimental approaches for disease-gene discovery such as transcriptomics and GWAS result in long, noisy, heterogeneous gene lists, which may be trimmed to a subset of likely relevant genes while leaving several false negatives. Computational gene classification approaches, especially those using genome-scale molecular interaction networks, are promising avenues for complementing such experimental findings by analytically expanding observed gene lists based on the functional relatedness between genes. We previously introduced the network-based gene classification approach, GenePlexus, which was rigorously benchmarked to show state-of-the-art performance, especially for predicting novel genes associated with biological processes and fine-grained phenotypes. Network-based gene classification performance,however, declines for diseases, especially when the inputs are omics and GWAS-based long gene lists. RESULTS: Here, we show that these disease gene lists span multiple biological processes spread across the molecular network, and we propose ModGenePlexus, a new network-based gene classification method that takes a two-stage approach. First, clustering and semi-supervised learning decomposes the input gene list into coherent, denoised network gene modules. Then, ModGenePlexus trains supervised (GenePlexus) classifiers for each module and aggregates predictions to return genome-wide rankings. We benchmarked ModGenePlexus across simulated data, transcriptomic signatures, and GWAS datasets (together spanning hundreds of diseases), showing improved recovery of known disease genes compared to GenePlexus. Beyond improved classification, the results of enrichment analysis of ModGenePlexus outputs are much more interpretable by virtue of revealing nuanced biological processes. Together, these results establish ModGenePlexus as a scalable, interpretable tool for gene classification of GWAS- and omics-derived gene lists across diverse biological contexts. AVAILABILITY AND IMPLEMENTATION: ModGenePlexus is freely available on GitHub at https://github.com/krishnanlab/ModGenePlexus, and the full source code and results supporting this study are available on Zenodo at https://zenodo.org/records/19857910.

Genome-Wide Association Study

Multipopulation GWAS for venous thromboembolism identifies novel loci followed by experimental validation in zebrafish.

Venous thromboembolisms (VTEs) are a leading cause of morbidity and mortality. Although many genetic risk factors have been identified, a substantial portion of the heritability remains unexplained. In this study, we employed a genome-wide association study (GWAS) for VTE across 9 international cohorts of the Global Biobank Meta-Analysis Initiative to address this question, along with in vivo functional validation. In this multipopulation GWAS (VTE cases, 27 987; controls, 1 035 290), 38 genome-wide significant loci were identified, 4 of which were potentially novel. For each autosomal locus, we performed gene prioritization using 7 independent, yet converging, lines of evidence. Through prioritization, we identified genes associated with VTE through GWAS and/or functional studies (eg, F5, F11, VWF, STAB2, PLCG2, TC2N), functionally validated those that did not have evidence other than GWAS (TC2N, TSPAN15), and discovered 1 not previously associated with coagulation (RASIP1). We evaluated the function of 6 prioritized genes with strong genetic evidence, including F7 as a positive control, using laser-mediated endothelial injury to induce thrombosis in zebrafish after CRISPR/Cas9 knockdown. From this assay, we have supportive evidence for the role of RASIP1 and TC2N in the modification of human VTE and suggestive evidence for STAB2 and TSPAN15. This study expands on the currently identified genomic architecture of VTE through biobank-based, multipopulation GWASs, in silico candidate gene predictions, and in vivo functional follow-up of candidate genes.

Zebrafish

Parietal Cortex Transcriptomics Refines Parkinson Disease GWAS Nomination and Highlights STAT3 as a Putative Upstream Glial Regulator.

Parkinson disease (PD) affects more than 1.1 million individuals in the United States and around 12 million worldwide. Although Genome Wide Association Studies (GWAS) have substantially advanced our understanding of PD genetic architecture, the regulatory mechanisms linking PD risk loci to disease-relevant gene expression remain incompletely characterized, limiting our ability to infer disease mechanisms from genetic associations. Here, we integrated disease-state parietal cortex transcriptomics with the International Parkinson's Disease Genomics Consortium (iPDGC) locus prioritization to refine PD gene nomination and identify biologically plausible candidates missed by GWAS-only approaches. Using bulk RNA-seq from 99 neuropathologically confirmed PD cases and 30 neuropathologically confirmed controls, we prioritized candidate genes across 78 loci and classified them according to concordance between genetic evidence and differential expression in diseased cortices. This integrative approach recovered candidate genes not captured by external GWAS-based prioritization methods and highlighted synaptic, lysosomal, and proteostasis pathways as major components of PD risk biology. Network and transcription factor analyses further suggested coordinated regulation of these genes, with STAT3 emerging as a putative upstream glial regulator. Together, these findings suggest that integrating disease-state transcriptomics with genetic prioritization can refine PD risk-gene nomination and uncover regulatory programs that may be missed by GWAS alone.

Journal Article

Integrative post-GWAS analysis prioritizes immune regulatory pathways and candidate effector signals in systemic lupus erythematosus.

BACKGROUND: Systemic lupus erythematosus (SLE) has a complex polygenic architecture, but translating genome-wide association signals into biologically interpretable candidates remains challenging. We applied an integrative post-GWAS framework to refine SLE-associated loci and prioritize candidate regulatory mechanisms. METHODS: European-ancestry SLE GWAS summary statistics from FinnGen and Bentham et al. were meta-analysed, comprising 8417 cases and 354,277 controls. After quality filtering, 6,782,131 SNPs were retained. Downstream analyses included LAVA regional prioritization, Bayesian colocalization with GTEx v8 whole-blood and spleen eQTLs, independent replication in the Julià et al. Spanish cohort, pathway enrichment, bivariate LAVA cross-trait local genetic correlation, and therapeutic annotation. RESULTS: The discovery meta-analysis identified 46 genome-wide significant SLE-associated loci, including putative novel signals requiring database/literature qualification. LAVA identified 14 candidate index variants across 12 high-confidence regions, of which nine index variants were retained as the primary prioritized set based on LAVA support and/or convergent regulatory evidence. The strongest association mapped to the chr6p21.3/MHC region (rs389884), where four genes showed colocalization support, including CLIC1 in whole blood and C4A in spleen. Because the chr6p21.3/MHC rs389884 region lead variant was unavailable for replication and no suitable proxy was identified, this signal was interpreted as an emerging candidate for functional validation rather than a replicated causal signal. Seven available variants replicated with concordant effects. An exploratory Roadmap immune chromatin-state overlap analysis placed 15 of 45 non-MHC lead variants (33.3%) directly, and 34 of 45 (75.6%) within ±10 kb, in active immune enhancer/promoter states. Pathway analyses highlighted type I interferon, JAK-STAT signaling, cytokine regulation, and antigen presentation, while bivariate LAVA analyses supported shared local genetic architecture with rheumatoid arthritis, systemic sclerosis, and Sjögren syndrome. CONCLUSIONS: This integrative post-GWAS analysis refines SLE association signals into biologically interpretable candidate regions and supports interferon and JAK-STAT signaling as central genetically supported pathways in SLE.

CLIC1

GWAS of Tau-Neurodegeneration Mismatch Identifies New Risk Loci for Susceptibility to Tau.

BACKGROUND: In Alzheimer's disease (AD), neurodegeneration is primarily attributed to the accumulation of tau neurofibrillary tangles. However, the distribution patterns of both tau pathology and neurodegeneration vary across different brain regions and among individuals. Moreover, multiple factors may influence the relationship between tau burden and neurodegenerative processes. Identifying the genetic architecture associated with deviation in the tau-neurodegeneration relationship can provide deeper mechanistic insights and guide the development of precision medicine strategies. METHODS: Here, I perform a genome-wide association study (GWAS) of cortical tau and thickness quantified by positron emission tomography (PET) and magnetic resonance imaging (MRI) in 794 participants from two cohorts of Alzheimer's disease Neuroimaging Initiative (ADNI) and A4. RESULTS: A GWAS was identified between the Tau/Neurodegeneration residual and two novel loci on chromosomes 7 and 14, with two SNPs (rs9323573 and rs9784993) exceeding the genome-wide significance threshold (p ≤ 5 × 10-8). SNP rs9323573 is located in STXBP6 on chromosome 14, while rs9784993 is in AKAP9 on chromosome 7, both of which were directly genotyped. The minor allele G of both SNPs (rs9323573, MAF = 0.221, p = 2.60 × 10-8; rs9784993, MAF = 0.197, p = 4.92 × 10-8) was associated with lower Tau/Neurodegeneration residuals, indicating higher-than-expected neurodegeneration given tau levels. CONCLUSION: GWAS of tau-related neurodegeneration identified two novel genetic variants in the loci AKAP9 and STXBP6 leading to higher than expected regional neurodegeneration given the tau level. Identifying genetic factors involved in tau-neurodegeneration mismatch may improve our understanding regarding the potential mechanistic downstream leading to susceptibility or resilience to tau pathology.

Humans

GWAS-based identification of a candidate gene and development of a predictive KASP marker for seed protein and oil contents in soybean.

BACKGROUND: Soybean [Glycine max (L.) Merrill] is one of the most widely cultivated crops worldwide. Its seeds contain about 40% protein and 20% oil, serving as essential nutrient sources for humans. Given the nutritional importance of seed protein and oil, identifying genes that regulate their levels is crucial for improving soybean seed quality. OBJECTIVE: This study aimed to identify genetic factors associated with seed protein and oil content using a genome-wide association study (GWAS). METHODS: Seed protein and oil contents were quantified in 192 soybean mutant accessions in a mutant diversity pool (MDP), and GWAS was conducted using 17,631 SNPs filtered from genotyping-by-sequencing. Expression of a candidate gene was examined across seed developmental stages (R5 to R7), and a significant SNP was converted into a Kompetitive Allele-Specific PCR (KASP) marker for validation. RESULTS: GWAS detected significant SNPs associated with seed protein and oil content. Chr20_7635098 was identified as a nonsynonymous SNP located in the exon of Glyma.20g042400. This gene showed differential expression across seed developmental stages between mutant accessions with contrasting protein and oil contents. The KASP marker for Chr20_7635098 was validated using the MDP and six domestic soybean cultivars showing predictive accuracies of ≥ 80.50% for protein content and ≥ 61.18% for oil content. CONCLUSION: Overall, this study identified a candidate gene linked to both seed protein and oil content, providing valuable insights for molecular breeding strategies aimed at efficiently improving these nutritional traits.

Glycine max

GWAS of CRP response to statins further supports the role of APOE in statin response: A GIST consortium study.

Statins are first-line treatments in the primary and secondary prevention of cardiovascular disease. Clinical studies show statins act independently of lipid-lowering mechanisms to decrease C-reactive protein (CRP), an inflammation marker. We aim to elucidate genetic loci associated with CRP statin response. CRP statin response is the change in log-CRP between off-treatment and on-treatment measurements. Cohort-level Genome-Wide Association Studies (GWAS) of CRP response were performed using 1000 Genomes imputed data, testing &#x223c;10 million common genetic variants. GWAS meta-analysis combined results from seven cohorts and clinical trials totalling 14,070 statin-treated individuals of European ancestry within the GIST consortium. Secondary analyses included statin-by-placebo interaction analyses, and lookups in African ancestry cohorts. Our GWAS identified two genome-wide significant (P&#x202f;<&#x202f;5e-8) loci: APOE and HNF1A for CRP statin response corrected for baseline CRP. The missense lead variant rs429358 at APOE, contributing to the APOE-E4 haplotype, is a risk locus for dyslipidaemia, Alzheimer's and coronary artery disease (CAD). The HNF1A locus is associated with diabetes, cholesterol levels, and CAD. Both loci are also associated with baseline CRP levels, and neither locus achieved a significant (P&#x202f;<&#x202f;0.05) result from the statin v. placebo interaction meta-analysis using randomized clinical trial data. However, the interaction result (P-int=0.09) for APOE was suggestive and possibly underpowered. The APOE-E4 signal may therefore be associated with both CRP and LDL-cholesterol statin response. Combined with suggestions in the literature that APOE also leads to differential statin benefit in Alzheimer's, the APOE locus warrants further investigation for potential genetic effects on healthcare with statin treatment.

Humans

Integrative haplotype and SNP-based GWAS supports the identification of stable genomic loci controlling yield-related traits in soybean.

Soybean yield is vulnerable to environmental variation, therefore, it is important to detect and implement stable genomic regions associated with yield-related traits in soybean breeding programs. In this study, SNP and haplotype-based GWAS were conducted to reveal important candidate genomic regions and putative candidate genes associated with soybean yield-related traits. This study demonstrates that the integration of haplotype and SNP-based GWAS could improve the detection of genomic regions associated with complex traits, enhance statistical power, and facilitate the identification of biologically relevant candidate genes. Ten stable haplotype blocks and six stable SNPs were detected based on the integration of haplotype and SNP-based GWAS, respectively. Furthermore, multiple candidate genes associated with the yield-related traits were identified. For instance, six genes were identified as transporters, including Glyma.15G092800, encoding serine-type endopeptidase activity, Glyma.15G203300 encoding a major facilitator superfamily (MFS) sugar transporter, Glyma.04G163000, transmembrane transporter, and Glyma.04G164100, leucine-rich repeat receptor-like protein kinase (LRR-RLK), as the most promising candidate genes. Additionally, three genes involved in signaling and pathways of various phytohormones can be promising candidates for increasing seed yield through improving plant architecture in soybean plants. The identified superior haplotypes with favourable alleles will be useful for marker-assisted selection in future breeding programs in soybean.

DArT markers

Multi-Ancestry Survival GWAS of Substance Use Initiation in the ABCD Study.

BACKGROUND: Substance use initiation in adolescence is influenced by both genetic and environmental factors; however, large-scale genetic studies often treat initiation as a binary outcome and underuse longitudinal timing information. METHODS: We conducted time-to-event (survival) genome-wide association analyses (GWAS) of initiation for four outcomes-alcohol, nicotine, cannabis, and any substance use-using longitudinal follow-up data from the Adolescent Brain Cognitive Development (ABCD) Study. We performed ancestry-stratified GWAS within European (EUR), African (AFR), and Hispanic (HISP) groups, applying consistent quality control and covariate adjustment. Summary statistics were harmonized across ancestries and meta-analyzed using inverse-variance weighted fixed-effects and DerSimonian-Laird random-effects models. We evaluated genomic inflation and heterogeneity (Cochran's Q and I 2), identified independent lead variants at genome-wide and suggestive significance thresholds, and assessed cross-trait overlap of associated loci. RESULTS: In the multi-ancestry meta-analysis, we observed suggestive association signals across traits (minimum p-values: alcohol ~ 1 &#xd7; 10-7, any ~ 1 &#xd7; 10-7, cannabis ~ 5 &#xd7; 10-8, nicotine ~ 1 &#xd7; 10-8). Nicotine initiation showed one genome-wide significant variant in both fixed- and random-effects meta-analyses (p < 5 &#xd7; 10-8). Across traits, suggestive loci demonstrated limited overlap, with the strongest concordance between alcohol and any substance use, consistent with shared liability. Heterogeneity statistics indicated that some loci exhibited cross-ancestry variation in effect estimates. CONCLUSIONS: Survival GWAS leveraging initiation timing can identify genetic signals that may be missed by binary designs and enables principled multi-ancestry synthesis. Our results highlight both shared and trait-specific genetic contributions to early substance initiation and provide a foundation for downstream functional annotation and integrative modeling with environmental risk factors. These findings demonstrate the value of incorporating developmental timing into genetic discovery and provide a framework for integrating longitudinal risk modeling with genomic analyses.

ABCD

Rethinking GWAS: how lessons from genetic screens and artificial intelligence could reveal biological mechanisms.

MOTIVATION: Modern single-cell omics data are key to unraveling the complex mechanisms underlying risk for complex diseases revealed by genome-wide association studies (GWAS). Phenotypic screens in model organisms have several important parallels to GWAS which the author explores in this essay. RESULTS: The author provides the historical context of such screens, comparing and contrasting similarities to association studies, and how these screens in model organisms can teach us what to look for. Then the author considers how the results of GWAS might be exhaustively interrogated to interpret the biological mechanisms underpinning disease processes. Finally, the author proposes a general framework for tackling this problem computationally, and explore the data, mechanisms, and technology (both existing and yet to be invented) that are necessary to complete the task. AVAILABILITY AND IMPLEMENTATION: There are no data or code associated with this article.

Genome-Wide Association Study

STABIX: summary-statistic-based GWAS indexing and compression.

MOTIVATION: Genome-wide association studies (GWAS) are widely used to investigate the role of genetics in disease traits, but the resulting file sizes from these studies are large, posing barriers to efficient storage, sharing, and querying. This issue is especially important for biobanks like the UK Biobank that publish GWAS for thousands of traits, increasing the volume of data that must be effectively managed. Current compression and query methods reduce file sizes and allow for quick genomic position-based queries but do not provide utility for quickly finding loci based on their summary statistics. For example, finding all SNVs in a particular p-value range would require decompressing and scanning the whole file. We propose a new tool, STABIX, which introduces summary-statistic-based queries and improves upon the standard bgzip compression and Tabix query tool in both compression ratio and decompression speed. RESULTS: When applied to 10 GWAS files from PanUKBB, STABIX created smaller compressed data and indices than Tabix for all files, where bgzip and tbi files were an average of 1.2 times the size of STABIX compressed files and indexes. In the same 10 files, STABIX per gene decompression was, on average 7&#xd7; faster than Tabix per gene decompression, and achieved faster per gene decompression times for over 99% of nearly 20,000 genes. AVAILABILITY AND IMPLEMENTATION: Software freely available for download at GitHub: https://github.com/kristen-schneider/stabix/.

Genome-Wide Association Study

GWAS for Periodontitis Phenotypes Using Multi-Ancestry All of Us Research Platform.

Periodontitis is a multifactorial inflammatory disease whose pathogenesis is associated with intricate interactions between genetic and environmental factors. Leveraging electronic health records data from the All of Us Research Program, we stratified periodontitis by clinically relevant dimensions: stage, grade, and extent. Based on these phenotypes, we performed a multi-ancestry genome-wide association study, focusing on predominant ancestry populations of African, European, and Admixed American. Our study cohort comprised 3,881 periodontitis patients and a control group of 10,760 patients with dental caries and without periodontitis. Ancestry-specific GWAS revealed significant genetic associations (P<5&#xd7;10-8) in periodontitis grade phenotypes at the LINC00294 and CLMN loci in the African ancestry population and also confirmed via the multi-ancestry meta-analysis. In addition, the XYLT1 locus emerged as a significant signal associated with periodontitis grade phenotype in the admixed American GWAS. Our GWAS comparing periodontitis to dental caries in the admixed American population identified several significant loci, including RABGAP1L, previously linked to immune regulation, DCHS2, a cadherin-related gene involved in bone mineralization and tissue morphogenesis, and OSTM1, known to be crucial for bone remodeling. The findings of our study highlight the potential of integrating EHR and genomic data from large-scale biobanks to achieve informative dental phenotyping, uncover novel molecular insights into periodontal disease, and personalize treatment approaches.

Journal Article

Integrating machine learning and GWAS for variant prioritization in the INCIPE cohort highlights ABC transporter genes in chronic kidney disease.

INTRODUCTION: Chronic kidney disease (CKD) is a major public health challenge, affecting approximately 674 million people worldwide and representing one of the fastest-growing causes of mortality. Since CKD is frequently asymptomatic in its early stages, the identification of novel genetic biomarkers may improve early detection and risk stratification. Genome-Wide Association Studies (GWAS) have identified numerous genetic loci associated with CKD and related traits; however, their performance is often limited in small and imbalanced cohorts, where reduced statistical power increases both false-positive and false-negative findings. Machine learning (ML) approaches can complement conventional GWAS by prioritizing biologically relevant genetic signals from high-dimensional genomic data. METHODS: In this study, we implemented a nested ensemble (NCBC) model composed of an undersampler and a CatBoostClassifier (CBC) to prioritize candidate genetic variants associated with CKD in the INCIPE cohort. Prioritized variants were functionally annotated and evaluated through enrichment analyses, GTEx gene expression profiling, and protein-protein interaction network analyses. Genes identified by the CKDGen Consortium were analysed as an external reference set and used to validate the biological relevance of the prioritized results. RESULTS: The NCBC model outperformed conventional ML classifiers, achieving a ROC AUC score of 87.77%, compared to 50%-53% for the other evaluated models. Among the prioritized genes, 56.25% showed protein-protein interactions with genes previously reported by the CKDGen Consortium, whereas only 1.9% of randomly generated gene sets showed interactions. DISCUSSION: Our study demonstrates that the NCBC model improves the prioritization of biologically plausible candidate variants in a small and imbalanced CKD cohort. Functional analyses suggested ABC transporter-related genes, including ABCA13, ABCA4, and ABCC4 genes, as promising candidate for future validation, with ABCA4 showing substantial expression in kidney tissues. Overall, these findings support the integration of ML with GWAS to prioritize candidate genes and investigate the genetic architecture of complex diseases.

SNP prioritization

Trio-based GWAS reveals loci associated with different forms of isolated cleft lip.

Orofacial clefts (OFCs) are the most common craniofacial birth defect and comprise a diverse group of traits with complex and heterogeneous etiologies. Genetic studies of OFCs typically approach this diversity by stratifying cases into broad diagnostic classes, including cleft lip (CL), cleft palate (CP), and cleft lip with palate (CLP). Although this strategy has yielded important insights into OFC risk, it ignores the phenotypic heterogeneity within each subtype. CL exhibits marked phenotypic variability, involving differences in alveolar involvement, laterality, and sidedness that may reflect distinct etiologies. Given this phenotypic diversity within CL, we assembled a multi-ancestry cohort of 837 nonsyndromic CL case-parent trios with whole-genome sequencing and detailed phenotyping. We performed genome-wide association scans (GWAS) via transmission disequilibrium tests for CL overall and for 14 CL subtypes defined by involvement of the alveolus (with and without), laterality (uni- and bilateral), and sidedness (left and right). We identified four genome-wide significant loci. Two loci, IRF6 and 8q24.21, were both detected in the overall CL GWAS. PLCB1/PLCB4 and MAFB were detected in GWASs of alveolar cleft involvement and CL left sidedness, respectively. These subtype-specific associations were followed by case-only comparisons that reflect the presence or absence of alveolus cleft or left-sided bias of CL to confirm the specificity of the association signal to the particular subtype. Our results provide evidence of within-class CL subtype-specific genetic links for loci previously discussed in the context of primary OFC classes and demonstrate the value of granular OFC subtype characterization to capture trait-specific associations.

Alveolus Cleft

Integrated GWAS and methylation analysis identify DNMT3A as an important regulator of growth in rabbits.

The parameters of individual growth curve can serve as pseudo-phenotype for genetic evaluation in livestock. In this study, we compared five nonlinear growth models using post-weaning body weights of 706 New Zealand White rabbits. Under the best-fitting model, two parameters of mature weight and maturity rate were subjected to GWAS through single-step genomic BLUP framework that integrated phenotypic records from non-genotyped animals with 41,359 SNPs genotyped in 198 individuals. Association analysis identified 147 relevant genomic regions, and also highlighted DNMT3A as a promising candidate gene for further functional investigation. siRNA-mediated knockdown of DNMT3A significantly impaired myoblast proliferation. Whole-genome bisulfite sequencing of DNMT3A-knockdown myoblasts identified 69,480 differentially methylated regions (DMRs). Integrative analyses revealed substantial overlap between DMR-associated genes and GWAS candidate genes, with significant enrichment in vitamin B6 and tyrosine metabolism pathways. These findings suggest that DNMT3A may regulate rabbit growth via mediating DNA methylation of downstream genes.

Animals