PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Annotation transfer for genomics: measuring functional divergence in multi-domain proteins.

Annotation transfer is a principal process in genome annotation. It involves "transferring" structural and functional annotation to uncharacterized open reading frames (ORFs) in a newly completed genome from experimentally characterized proteins similar in sequence. To prevent errors in genome annotation, it is important that this process be robust and statistically well-characterized, especially with regard to how it depends on the degree of sequence similarity. Previously, we and others have analyzed annotation transfer in single-domain proteins. Multi-domain proteins, which make up the bulk of the ORFs in eukaryotic genomes, present more complex issues in functional conservation. Here we present a large-scale survey of annotation transfer in these proteins, using scop superfamilies to define domain folds and a thesaurus based on SWISS-PROT keywords to define functional categories. Our survey reveals that multi-domain proteins have significantly less functional conservation than single-domain ones, except when they share the exact same combination of domain folds. In particular, we find that for multi-domain proteins, approximate function can be accurately transferred with only 35% certainty for pairs of proteins sharing one structural superfamily. In contrast, this value is 67% for pairs of single-domain proteins sharing the same structural superfamily. On the other hand, if two multi-domain proteins contain the same combination of two structural superfamilies the probability of their sharing the same function increases to 80% in the case of complete coverage along the full length of both proteins, this value increases further to > 90%. Moreover, we found that only 70 of the current total of 455 structural superfamilies are found in both single and multi-domain proteins and only 14 of these were associated with the same function in both categories of proteins. We also investigated the degree to which function could be transferred between pairs of multi-domain proteins with respect to the degree of sequence similarity between them, finding that functional divergence at a given amount of sequence similarity is always about two-fold greater for pairs of multi-domain proteins (sharing similarity over a single domain) in comparison to pairs of single-domain ones, though the overall shape of the relationship is quite similar. Further information is available at http://partslist.org/func or http://bioinfo.mbb.yale.edu/partslist/func.

Computational Biology↗

Genomes of the ex-type strains of Elsinoë mangiferae and E. perseae, the causal agents of scab on mango and avocado.

Elsinoë species are slow-growing, hemibiotrophic to necrotrophic fungi that cause scab diseases on economically important fruit crops. Genome resources for many host-specific species remain limited. We report high-quality draft genome assemblies for the ex-type strains of Elsinoë mangiferae (CBS 226.50) and E. perseae (CBS 406.34), causal agents of mango and avocado scab, respectively. Among 5 approaches tested, a Nanopore-only NextDenovo assembly produced the most contiguous genomes, yielding 24.5 Mb (E. mangiferae) and 25.1 Mb (E. perseae) assemblies with 13 and 18 contigs, respectively, BUSCO completeness scores of ∼94%, and multiple putative telomere-to-telomere chromosomes. Gene prediction identified 9,134 and 9,243 genes, respectively. Functional annotation revealed enrichment of metabolic and regulatory pathways, including those involved in posttranslational modification, protein transport, and secondary metabolism. Carbohydrate-active enzyme repertoires were small but conserved, consistent with stealth pathogenicity strategies and low plant cell wall degradation. Both genomes encoded large secretomes (>850 proteins), diverse protease repertoires (>300 proteins), Ecp2-like effector proteins, and multiple biosynthetic gene clusters, including clusters with similarity to those associated with elsinochrome and ACT-toxin II biosynthesis, some of which may contribute to host-pathogen interactions and disease development. A large fraction of genes lacked functional characterization, suggesting incomplete databases and/or the presence of lineage-specific genes potentially involved in virulence or host adaptation. These genome resources fill critical gaps for underrepresented Elsinoë species and provide taxonomically anchored references essential for diagnostics, comparative genomics, and research into the molecular basis of host specificity and pathogenicity in scab-causing fungi.

Persea↗

B cell pathways implicate shared genetic architecture between schizophrenia and immune-mediated diseases.

BACKGROUND: Schizophrenia and immune-mediated diseases are globally prevalent and highly heritable conditions that frequently co-occur, posing major public health burdens. However, their shared genetic architecture remains poorly understood. METHODS: We applied the bivariate causal mixture model (MiXeR) to investigate the polygenic overlap between schizophrenia and eight common immune-mediated diseases, using genome-wide association study summary statistics comprising 2,489 to 67,323 cases and 9,066 to 497,622 controls. Shared loci were identified through conditional/conjunctional false discovery rate (cond/conjFDR), local genetic correlation (LAVA), and colocalization analyses. Subsequently, gene mapping, functional annotation, expression-trait association, and drug-gene interaction analyses were performed to explore shared genes and enriched pathways, and genetic risk scores (GRS) from the UK Biobank were used to validate the findings. RESULTS: MiXeR estimated substantial polygenic overlap between schizophrenia and immune-mediated diseases, and conjFDR identified 133 shared loci, with eight prioritized through local genetic correlation and colocalization signals. These eight loci were mapped to 85 protein-coding genes enriched in pathways essential for B cell function. Among them, S-PrediXcan analyses identified 14 genes whose expression in brain tissues or blood was associated with both diseases. These genes also interact with immunomodulatory or antihypertensive drugs. Additionally, 11 of the 14 genes were linked to innate immunity and/or cognitive traits. Using UK Biobank data, we further confirmed that overall, shared gene, and B cell activation and receptor signaling pathway–specific genetic risk for schizophrenia is associated with immune-mediated disease susceptibility. CONCLUSIONS: These findings underscore the shared genetic architecture of schizophrenia and immune-mediated diseases, advancing insights at the interface of psychiatric genetics and immunology.

Schizophrenia↗

Identifying fundamental gaps in functional metagenomics: a step towards unlocking microbiome research potential.

Incomplete functional annotation limits biological interpretation in microbiome studies and their translational potential. Poor annotation arises from multiple causes, with incomplete gene-protein-reaction mapping being one tractable yet under-examined contributor. We address this gap by developing a comprehensive hierarchical framework that systematically integrates gene families in UniRef, proteins in UniProt, and metabolic reactions in MetaCyc and BioCyc through UniProtKB accession, EC number, and Pfam-domain matching. Applied to a human gut metagenome dataset via HUMAnN3, our MetaCyc-based mapping recovers up to 2.3-fold more unique reaction identifiers than the default pipeline and increases reaction prevalence across samples from ≈32% to 52% core reactions, addressing the data sparsity that limits statistical and machine-learning applications in microbiome research. Biological plausibility for the tested functions was supported by positive and negative controls: gut-microbial hormone-metabolism reactions previously linked to this dataset were recovered, while vertebrate-specific hormone-metabolism reactions remained correctly undetected. These gains derive from systematic database integration alone, without predictive algorithms, indicating that a tractable, mapping-related component of functional dark matter and data sparsity in microbiome studies is directly addressable. Because Pfam- and BioCyc-derived mappings trade specificity for coverage, confidence in any individual reaction assignment depends on the supporting evidence tier and source database.

Humans↗

Chromosomal-level genome assembly of minute pirate bug Orius nagaii Yasunaga, 1993 (Hemiptera: Anthocoridae).

Species of the genus Orius, diminutive predatory insects that act as natural enemies of other arthropods, are frequently employed in agricultural pest management for controlling various pests, such as thrips, mites, aphids, whiteflies, etc. However, the scarcity of high-quality genomic resources for these predators hinders our comprehension of their population evolution and predation ecology. Consequently, we assembled and annotated a chromosomal-scale genome of Orius nagaii by collating PacBio and Illumina sequencing and Hi-C genomic analysis techniques. The final genome assembly size 152.62 Mb, with scaffold and contig N50 lengths of 11.53 and 2.39 Mb, respectively. It is organized into 12 pairs of autosomes and a pair of XY sex chromosomes. The quality assessment of the genomic data with BUSCO revealed a completeness of 98.5% (n = 1,367). Also, 11,917 protein-coding genes were discovered, with 94.28% of them having functional annotations. The high-quality genome of O. nagaii produced serves as a valuable resource for comprehending the interactions between predatory natural enemies and hosts, along with their evolutionary trajectories.

Animals↗

EucaMOD: a comprehensive multi-omics database for functional genomics research and molecular breeding of fast-growing eucalyptus trees.

Eucalyptus, one of the most widely planted plantation tree species globally, is primarily found in tropical and subtropical regions and contributes significantly to economic and social benefits. With advances in sequencing technologies, there is an increasing demand for the systematic analysis of multi-omics data among Eucalyptus species to enhance genetic breeding efforts. Although several early genomic databases have been established for eucalyptus, they have not been updated in a timely manner and lack recent multi-omics data, rendering them insufficient for current research needs. To address this gap, we developed the eucalyptus multi-omics database (EucaMOD, http://eucalyptusggd.net/eucamod), a comprehensive resource for cross-omics studies. In this study, we functionally annotated 45 eucalyptus genomes and structurally annotated 15, conducting comparative genomics and pan-proteomics analyses across all genomes. Additionally, we analyzed eucalyptus transcriptome, epigenome, and variome data through standardized workflows, enabling the in-depth mining and reanalysis of multi-omics datasets. EucaMOD is the most comprehensive multi-omics database for eucalyptus to date and includes data from 45 genomes (39 species), 870 mRNA-seq samples, 17 miRNA-seq samples, 52 epigenomic datasets (histone modifications and transcription factor binding), and genetic variation data from 1219 samples. To support functional genomics and molecular breeding research, the database is organized into the following 11 modules: Home, Species, Genomics, Comparative genomics, Pan-proteomics, Transcriptomics, Epigenetics, Variomics, Tools, Download, and Help. EucaMOD also offers online analysis tools for data mining, providing free public services to aid eucalyptus gene function and genetic engineering studies.

Eucalyptus↗

Microbial and functional shifts between flare and remission in a single-center cohort of children with inflammatory bowel disease.

BACKGROUND: Gut microbial dysbiosis is central to the pathogenesis of inflammatory bowel disease (IBD). While gut microbiome differences between patients with and without IBD are well established, microbiome changes associated with disease activity and remission remain limited, particularly in paediatric populations. AIM: To examine intra-individual taxonomic and functional gut microbiome changes during transition from active flare to remission under maintenance immunosuppression in a pilot single-center Singapore cohort of children with IBD. METHODS: Paired stool samples and clinical data were collected from seven patients with paediatric IBD [5 Crohn's disease (CD), 2 ulcerative colitis; &#x2264; 18 years] during active disease/flare (visit 1; Pediatric CD Activity Index/Pediatric Ulcerative Colitis Activity Index &#x2265; 10) and subsequent clinical remission (visit 2; Pediatric CD Activity Index/Pediatric Ulcerative Colitis Activity Index < 10). Samples underwent shotgun metagenomic sequencing for high-resolution taxonomic profiling and functional annotation of Kyoto Encyclopaedia of Genes and Genomes pathways. RESULTS: Gut microbial diversity was reduced during flare compared to remission, with Actinobacteria abundance significantly higher in remission. Two distinct microbial clusters differentiated flare and remission states: The remission cluster was enriched with Bifidobacterium adolescentis, Bifidobacterium dentium, Lactobacillus gasseri, Faecalibacterium prausnitzii, while the flare state showed increased Klebsiella pneumoniae. Remission was further characterized by a downregulation of pathogenic microbes and an upregulation of beneficial microbes including a higher abundance of the butyrate producer Anaerostipes hadrus (P = 0.046). Microbial functional genes enriched in remission were predominantly associated with metabolic pathways including vitamin and cofactor biosynthesis, as well as carbohydrate, amino acid, and lipid metabolism. CONCLUSION: The transition from flare to remission in Singaporean children with IBD is characterized by functional remodeling of the gut microbiome, which may contribute to recovery processes related to intestinal barrier integrity, cellular maintenance, and tissue repair. Targeted modulation of the gut microbiome may help sustain remission in paediatric IBD.

Functional shift↗

Application of a Translational Research Platform to Unveil Efficacy Signals and Mechanisms of Resistance of FGFR Inhibitors in Multiple FGFR-Altered Solid Tumors.

PURPOSE: The predictive value of fibroblast growth factor receptor (FGFR) amplifications (amp) and the role of FGFR mutations (mut) beyond known activating variants remain unclear. We aimed to establish a translational research platform to characterize FGFR alterations (alt) and explore their potential as predictive biomarkers for FGFR-targeted agents. EXPERIMENTAL DESIGN: This ambispective study included a retrospective analysis of patients with FGFR-alt tumors treated with selective FGFR inhibitors (FGFRi) and a prospective collection of longitudinal tumor samples. Patient-derived xenografts (PDX) were generated to investigate FGFRi mechanisms of action and resistance. Molecular characterization included genomic, transcriptomic, proteomic, and functional analyses using the Functional Annotation for Cancer Treatment (FACT) assay. RESULTS: Among 36 retrospectively analyzed patients, clinical benefit from FGFRis was observed in cases with FGFR mRNA overexpression or FGFR2/11q co-amp, but no association was found with the amplification levels. In archival tumor samples, exploratory proteomic analysis showed FGFR1-4 protein expression in 78% of FGFR1/2-amp tumors detected by fluorescence in situ hybridization. RNA sequencing identified a higher prevalence of FGFR mRNA overexpression than proteomic analysis. Among patients harboring FGFR-mut, only one bladder cancer with an FGFR3-mut S249C derived benefit. FACT assay supported the functional activity of selected variants, including FGFR3 T689M, and suggested potential resistance mechanisms involving PI3K/PTEN and MAPK pathway co-alterations. A prospective FGFR-alt PDX biorepository enabled exploratory biomarker analyses, supporting the hypothesis that FGFR1-4 mRNA expression may better reflect FGFR dependency than genomic alterations alone. CONCLUSIONS: These findings highlight the complexity of FGFR-driven oncogenesis and support integrative molecular approaches to refine patient selection for FGFR-targeted therapies.

Humans↗

Chromosome-level genome assembly and annotation of the porcupine fish (Diodon hystrix).

The porcupinefish (Diodon hystrix), a coral reef teleost, is widely distributed in tropical/subtropical waters of the Pacific, Atlantic, Indian Oceans, and Mediterranean Sea. It shares easily recognizable features with pufferfish, such as body inflation and spines. Additionally, its culinary value makes D. hystrix a highly desirable species in many tropical coastal regions, with considerable market potential. However, lack of a high-quality genome hindered further studies on its reproduction, molecular biology, and genomic improvement. Here, we assembled the chromosome-scale genome using PacBio HiFi, ultra-long reads, and Hi-C. Of the 713.62&#x2009;Mb genome, 98.63% anchored to 23 chromosomes (scaffold N50: 31.52&#x2009;Mb) with 39.82% repetitive sequences. The assembled genome achieved a BUSCO completeness score of 97.7%, with 23,171 protein-coding genes predicted, 22,221 of which were functionally annotated. Phylogenetic analysis identified D. hystrix's evolutionary relationships with other species in the Tetraodontiformes. In summary, the high-quality genome of D. hystrix sheds light on valuable insights into genome size evolution, and provides a valuable resource for exploiting genomic study and breeding applications in this species.

Animals↗

Chromosome-level genome assembly and annotation of Spinibarbus caldwelli.

Spinibarbus caldwelli is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in southern China, the lack of genomic resources has hampered the genetic breeding and conservation. Here, we release a chromosome-level genome assembly for S. caldwelli using PacBio HiFi long-reads, Illumina short-reads, and Hi-C sequencing data. The final genome assembly is 1.77&#x2009;Gb in size, with a contig N50 of 24.27&#x2009;Mb. Using Hi-C scaffolding, 99.14% of the contigs were successfully anchored to 50 chromosomes, resulting in a scaffold N50 of 35.29&#x2009;Mb. The final genome assembly shows a BUSCO completeness of 98.27%. The assembled genome contains 49.41% repetitive sequences and 51,505 predicted genes, 90.83% of which have been functionally annotated. This genome provides a genetic basis for S. caldwelli, facilitating the exploration of Cyprinid phylogeny, genetic improvement, and conservation efforts.

Animals↗

Transcriptomic insights into the coordinated regulation of signaling, apoptosis, immunity, and metabolism during Sinonovacula constricta larval metamorphosis.

Metamorphosis is a critical ontogenetic transition for marine bivalves, marking the shift from planktonic to benthic lifestyles, where successful transformation dictates survival. The razor clam Sinonovacula constricta is economically important; however, low larval metamorphosis rates remain a major bottleneck in seedling production. To elucidate the mechanisms governing this process, we performed a comparative transcriptome analysis of S. constricta larvae at pre- and post-metamorphosis stages using Illumina sequencing. A total of 3701 differentially expressed genes (DEGs) were identified, including 3254 up-regulated and 447 down-regulated genes. Functional annotation of the respective top 20 significantly up-regulated and down-regulated DEGs indicated their potential pivotal roles in signal transduction (e.g., up-regulated: CAV1, CHRNA2; down-regulated: APP, NOTCH1), cellular proliferation and differentiation (e.g., up-regulated: TUBA, EGF1; down-regulated: KIF23, TTC25), transcriptional and epigenetic regulation (e.g., up-regulated: NFIL3; down-regulated: OVO, HMX1), substance transport (e.g., up-regulated: LRP2, LRP1B; down-regulated: SLC51A, Slc33a1), substance metabolism (e.g., up-regulated: CPK3, CYP26A1; down-regulated: RDMT1, ADAC), immunomodulation (e.g., up-regulated: CPN2, CRISP2), and protein homeostasis (e.g., up-regulated: HSP27, NAS-27). Functional enrichment analysis further revealed that DEGs were significantly enriched in pathways related to signal transduction and developmental regulation (e.g., Ras, TNF), cell death and homeostasis (e.g., apoptosis), immune responses (e.g., Toll-like receptor), energy metabolism (e.g., lipid), cardiovascular related (e.g., Fluid shear stress), cell junction and architecture (e.g., Tight junction), and infectious disease (e.g., measles). These results suggest a synergistic interplay between signaling, apoptosis, immunity, and metabolism during S. constricta metamorphosis. This study advances our understanding of marine bivalve metamorphosis and offers candidate genes for further mechanistic studies.

Animals↗

Meta-QTL Analysis Reveals Consensus Genomic Regions and Candidate Genes for Resistance to Sudden Death Syndrome in Soybean.

Sudden death syndrome (SDS), caused by Fusarium virguliforme, is one of the most economically important diseases limiting soybean production worldwide. Although numerous quantitative trait loci (QTL) associated with SDS resistance have been reported, inconsistencies among mapping populations, marker systems, and experimental conditions have hindered the identification of robust resistance loci for soybean improvement. In this study, a comprehensive meta-analysis was conducted to integrate published QTL and identify stable consensus genomic regions associated with SDS resistance. After a systematic literature survey and data curation, 153 QTL derived from 14 linkage-mapping studies were analyzed using a custom R-based workflow, resulting in the identification of 23 consensus meta-QTL (MQTL) distributed across 17 chromosomes. Several MQTL, particularly those located on chromosomes 6, 8, 18, and 20, were supported by multiple independent studies and represented major genomic hotspots for SDS resistance. Physical localization and functional annotation of these MQTL identified 217 candidate genes, including genes predicted to be involved in plant defense, signal transduction, transcriptional regulation, and secondary metabolism. Gene Ontology enrichment analysis identified response to salicylic acid as the only biological process that remained significant after FDR correction, whereas Kyoto Encyclopedia of Genes and Genomes pathway analysis did not identify significantly enriched pathways. Independent support using five published genome-wide association studies further supported several MQTL, especially those on chromosomes 6, 18, and 20, thereby increasing confidence in these genomic regions. The identified MQTL and prioritized candidate genes provide potential genomic resources for future marker development, improvement applications, and functional validation aimed at improving soybean resistance to SDS.

Fusarium virguliforme↗

MetagenomicKG: a knowledge graph for metagenomic applications.

MOTIVATION: The sheer volume and variety of genomic content within microbial communities makes metagenomics a field rich in biomedical knowledge. To traverse these complex communities and their vast unknowns, metagenomic studies often depend on distinct reference databases, such as the Genome Taxonomy Database (GTDB), the Kyoto Encyclopedia of Genes and Genomes (KEGG), and the Bacterial and Viral Bioinformatics Resource Center (BV-BRC), for various analytical purposes. These databases are crucial for the genetic and functional annotation of microbial communities. Nevertheless, the inconsistent nomenclature or identifiers of these databases present challenges for effective integration, representation, and utilization. Knowledge graphs (KGs) offer an appropriate solution by organizing biological entities from different databases to standardized identifiers, allowing their interrelations to be captured into a cohesive network regardless of the naming conventions used in each source. The graph structure not only facilitates the unveiling of hidden patterns but also enriches our biological understanding with deeper insights. Despite KGs having shown potential in various biomedical fields, their application in metagenomics remains underexplored. RESULTS: We present MetagenomicKG, a novel knowledge graph specifically tailored for metagenomic analysis. MetagenomicKG integrates taxonomic, functional, and pathogenesis-related information on the human microbiome sourced from various databases, and further connects these with existing biomedical KGs to expand the biological network. Through various case studies involving the human microbiome, we demonstrate its utility in enabling hypothesis generation regarding the relationships between microbes and diseases, generating sample-specific graph embeddings, and providing robust pathogen prediction. CODE AVAILABILITY: The source code and technical details for constructing the MetagenomicKG and reproducing all analyses are available on GitHub at https://github.com/KoslickiLab/MetagenomicKG. The data used in this manuscript, including the pre-built files and use case input data, are archived on Zenodo with DOI: 10.5281/zenodo.17546861.

Metagenomics↗

Whole genome sequence-based association analysis of African American individuals with bipolar disorder and schizophrenia.

In studies of individuals of primarily European genetic ancestry, common and low-frequency variants and rare coding variants have been found to be associated with the risk of bipolar disorder (BD) and schizophrenia (SZ). However, less is known for individuals of other genetic ancestries or the role of rare non-coding variants in BD and SZ risk. We performed whole genome sequencing of African American individuals: 1,598 with BD, 3,295 with SZ, and 2,651 unaffected controls (InPSYght study). We increased power by incorporating 14,812 jointly called psychiatrically unscreened ancestry-matched controls from the Trans-Omics for Precision Medicine (TOPMed) Program for a total of 17,463 controls. To identify variants and sets of variants associated with BD and/or SZ, we performed single-variant tests, gene-based tests for singleton protein truncating variants, and rare and low-frequency variant annotation-based tests with conservation and universal chromatin states and sliding windows. We found suggestive evidence of BD association with single-variants on chromosome 18 and of lower BD risk associated with rare and low-frequency variants on chromosome 11 in a region with multiple BD GWAS loci, using a sliding window approach. We also found that chromatin and conservation state tests can be used to detect differential calling of variants in controls sequenced at different centers and to assess the effectiveness of sequencing metric covariate adjustments. Our findings reinforce the need for continued whole genome sequencing in additional samples of African American individuals and more comprehensive functional annotation of non-coding variants.

Journal Article↗

A high-quality chromosome-level genome assembly and annotation of the giant freshwater prawn (Macrobrachium rosenbergii).

The giant freshwater prawn, Macrobrachium rosenbergii, is native to Southeast Asia and is used in aquacultural practices worldwide. It is considered advantageous because of its rapid growth, high nutritional value, and economic benefits. As one of the three major freshwater aquaculture shrimp sources in China, a high-quality genome resource is of great significance for promoting the germplasm improvement of varieties. This study presents a high-quality chromosome-level genome assembly of M. rosenbergii that was generated by combining PacBio, MGI, and Hi-C reads. The assembled genome was 2.96&#x2009;Gb in size, with a contig N50 of 0.64&#x2009;Mb and a scaffold N50 of 55.76&#x2009;Mb, which was positioned on 59 pseudo-chromosomes. The Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis for genome assembly reached 94.37%. In total, 27,111 protein-coding genes were identified, of which 25,470 were functionally annotated. These results provide a foundation for future research into adaptive evolution, genomics, and molecular breeding in M. rosenbergii.

Animals↗

Integrating machine learning and GWAS for variant prioritization in the INCIPE cohort highlights ABC transporter genes in chronic kidney disease.

INTRODUCTION: Chronic kidney disease (CKD) is a major public health challenge, affecting approximately 674 million people worldwide and representing one of the fastest-growing causes of mortality. Since CKD is frequently asymptomatic in its early stages, the identification of novel genetic biomarkers may improve early detection and risk stratification. Genome-Wide Association Studies (GWAS) have identified numerous genetic loci associated with CKD and related traits; however, their performance is often limited in small and imbalanced cohorts, where reduced statistical power increases both false-positive and false-negative findings. Machine learning (ML) approaches can complement conventional GWAS by prioritizing biologically relevant genetic signals from high-dimensional genomic data. METHODS: In this study, we implemented a nested ensemble (NCBC) model composed of an undersampler and a CatBoostClassifier (CBC) to prioritize candidate genetic variants associated with CKD in the INCIPE cohort. Prioritized variants were functionally annotated and evaluated through enrichment analyses, GTEx gene expression profiling, and protein-protein interaction network analyses. Genes identified by the CKDGen Consortium were analysed as an external reference set and used to validate the biological relevance of the prioritized results. RESULTS: The NCBC model outperformed conventional ML classifiers, achieving a ROC AUC score of 87.77%, compared to 50%-53% for the other evaluated models. Among the prioritized genes, 56.25% showed protein-protein interactions with genes previously reported by the CKDGen Consortium, whereas only 1.9% of randomly generated gene sets showed interactions. DISCUSSION: Our study demonstrates that the NCBC model improves the prioritization of biologically plausible candidate variants in a small and imbalanced CKD cohort. Functional analyses suggested ABC transporter-related genes, including ABCA13, ABCA4, and ABCC4 genes, as promising candidate for future validation, with ABCA4 showing substantial expression in kidney tissues. Overall, these findings support the integration of ML with GWAS to prioritize candidate genes and investigate the genetic architecture of complex diseases.

SNP prioritization↗

Unraveling 'F' factor: towards a genetic-clinical framework for the musculoskeletal-heart crosstalk in metabolic aging.

BACKGROUND: The rising co-occurrence of cardiometabolic diseases and musculoskeletal degeneration poses a critical challenge to healthy aging, yet the shared biological mechanisms underlying this multimorbidity remain poorly defined. This study aimed to establish an integrative clinical-genetic framework to elucidate the common frailty factor, the 'F' factor, that captures the systemic vulnerability linking cardiometabolic multimorbidity (CMM) and musculoskeletal aging. METHODS: Utilizing the prospective China Health and Retirement Longitudinal Study (CHARLS) cohort, we developed and validated novel Frailty-Integrated Indices for CMM risk prediction, evaluated with machine learning models interpreted via SHapley Additive exPlanations (SHAP). Independently, we applied genomic structural equation modeling (Genomic-SEM) to integrate genome-wide association data from six traits-coronary artery disease, type 2 diabetes, hypertension, bone mineral density, frailty, and telomere length-to model a shared latent genetic factor ('F' factor). This was followed by multivariate GWAS, fine-mapping, transcriptome-wide association study (TWAS), gene-based analysis, and functional annotation to prioritize causal genes, pathways, and cell types. RESULTS: Clinically, several Frailty-Integrated Indices significantly improved CMM risk prediction, with the optimal model achieving an AUC of 0.727. Genetically, we modeled a significant shared latent genetic factor ('F' factor), pinpointing novel risk loci and implicating key genes such as APOE and SLC22A3. These genes were enriched in pathways including cellular senescence and cholesterol metabolism and showed specific expression patterns in developmental brain stages and across multi-organ endothelial cells. CONCLUSION: Our findings provide converging evidence for Musculoskeletal&#x2011;Heart crosstalk of metabolic aging and inferred the 'F' factor as a genetic correlate of a transdiagnostic state, which links genetic predisposition to metabolic dysregulation, and systemic functional decline. This work provides a multi-level biological characterization of multimorbidity liability, informing early-risk detection and preventive strategies for complex aging-related comorbidities.

Humans↗

A sequence-based classifier distinguishes phenotype-associated genes from other gene models in plants.

Only a small fraction of annotated plant genes possess experimentally validated associations with specific phenotypes. Phenotype-associated genes have distinct structural, molecular, and evolutionary characteristics compared with nonvalidated gene models. Here, we develop a simple classifier that uses sequence and evolutionary features, which can be generated for any species with an annotated reference genome assembly, to accurately distinguish phenotype-associated genes from both the overall population of annotated gene models and a specific set of genes identified as being tolerant of premature stop mutations. A model trained solely on genes from maize (Zea mays) identifies and prioritizes rice (Oryza sativa) and Arabidopsis (Arabidopsis thaliana) genes that are highly enriched in genes with experimentally validated links to phenotypes in both of these evolutionarily distant species. Gene models predicted to have a higher probability of being linked to phenotypes display patterns consistent with known biological properties of phenotype-associated genes. Notably, the sets of genes predicted to have a high probability of being linked to phenotype variation do not consist exclusively of well-characterized gene families but included many uncharacterized gene families carrying domains of unknown function. The quantitative scores generated by this model offer a valuable resource for prioritizing and exploring the vast number of uncharacterized gene models in plants, reducing the risk of failure in future reverse genetic efforts and potentially accelerating gene discovery and functional annotation in crops.

Phenotype↗