PubMed HealthSearch

SEARCH · PubMed Health

Results for “data curation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa

Omics in hereditary optic neuropathies: A systematic review of clinical studies with an integrated point of view.

Hereditary optic neuropathies are characterized by bilateral visual loss due to the degeneration of retinal ganglion cells, resulting in optic nerve degeneration and atrophy. Although the genetic origin of the main isolated and syndromic hereditary optic neuropathies has been characterized, the clinical phenotypes exhibit significant and poorly understood variability in both penetrance and expressivity. Additionally, the genetic and environmental factors that influence the onset of these optic neuropathies remain poorly understood, with limited biomarkers to predict disease progression or as readouts for therapeutic trials. Data-driven omics strategies allow deep phenotyping to improve our understanding of pathophysiological mechanisms and to search for new biomarkers and therapeutic targets. We explore whether the omics strategies applied to patients with hereditary optic neuropathies have provided such new insights. MEDLINE, Web of Science and EMBASE databases were screened for studies with terms relating to hereditary optic neuropathies, transcriptomics, epigenomics, proteomics, metabolomics and lipidomics in clinical studies exploring patients' samples. Out of 1244 references identified, 22 articles were included after double-masked data curation. These articles focused only on the 3 main forms of hereditary optic neuropathies, namely, OPA1-related dominant optic atrophy (n = 4), Leber hereditary optic neuropathy (n = 13), and Wolfram syndrome (n = 5). While the methodological designs and results of these studies were highly heterogeneous, they revealed molecular alterations that we have attempted to discuss at the integrated multi-omics level. This data integration highlighted several common pathophysiological mechanisms such as energetic impairment, endoplasmic reticulum stress, proteotoxic and oxidative stresses, lipid remodeling and altered amino acid and purine metabolisms, while suggesting potential new biomarkers and therapeutic targets. These findings underscore the potential of integrated multi-omics approaches to deepen our understanding of the phenotypic complexity of hereditary optic neuropathies and to support the development of innovative diagnostic and therapeutic strategies.

Humans

Unveiling non-small cell lung cancer treatment effect heterogeneity: a comparative analysis of statistical methods.

BACKGROUND: For patients with advanced non-small cell lung cancer lacking targetable genomic alterations, the impact of clinicogenomic characteristics on the effectiveness of combining chemotherapy with immunotherapy is unclear. METHODS: We evaluated 4 statistical methods for detecting heterogeneous treatment effects related to clinical factors, including programmed death-ligand 1 expression, tumor mutation burden, and stage at diagnosis, using the American Association for Cancer Research Project Genomics Evidence Neoplasia Exchange BioPharma Collaborative dataset supplemented with institutional data collected under the same data curation model. A 2-sided P value of no more than .05 was used to denote statistical significance for all analyses. RESULTS: The mixture model revealed 2 latent subgroups: in one subgroup, there was no meaningful treatment effect, with average progression-free survival (PFS) only 5% longer with immunotherapy alone (95% confidence interval [CI] = -19% to 35%); in the second subgroup, immunotherapy alone was associated with a 35% decrease in average PFS (95% CI = -59% to 2%), corresponding to a ratio in treatment effects of 1.62 (95% CI = 1.02 to 2.57). There was a marginal association between lower tumor mutation burden levels and membership in the subgroup with improved PFS following receipt of chemoimmunotherapy. The causal survival forest highlighted the importance of tumor mutation burden (variable importance ranking: 1) and programmed death-ligand 1 (variable importance ranking: 3) when assessing heterogeneity. In contrast, the accelerated failure time and Cox proportional hazards models did not detect any statistically significant heterogeneous treatment effects. In simulations, the mixture model identified heterogeneous treatment effects more frequently than other methods, especially with weak covariate relationships, demonstrating its utility for informing personalized treatment approaches. CONCLUSIONS: The application of novel statistical methods to large scale clinico-genomic databases offers an opportunity to more accurately identify heterogeneous treatment effects in some settings as compared to traditional statistical methods. Applying such methods to the AACR Project GENIE BPC non-small cell lung cancer data indicated a potential association between decreasing tumor mutation burden and improved outcomes with chemoimmunotherapy as compared to immunotherapy alone.

Humans

Mitogenomic and phylogenomic analyses identify a cohesive Western Atlantic lineage within the Narcine complex (Torpediniformes: Narcinidae).

BACKGROUND: Accurate species delimitation within electric rays of the genus Narcine has been hindered by overlapping morphological characters and limited molecular resolution in previous single-locus studies. This study aims to evaluate phylogenetic relationships and species boundaries within the Narcine species complex across the Western Atlantic using complete mitochondrial genomes. METHODS AND RESULTS: Seven complete mitogenomes were newly assembled from individuals representing distinct morphotypes sampled across geographically widespread Western Atlantic localities and analyzed together with publicly available reference sequences. Mitochondrial protein-coding genes (PCGs) were examined using concatenated nucleotide and amino acid datasets under partitioned maximum-likelihood frameworks. Both approaches recovered highly congruent topologies, consistently supporting a single, well-defined western Atlantic mitochondrial lineage with low internal divergence (0.04-2.13%). Species delimitation analyses based on multiple methods yielded partially congruent results but consistently identified a dominant lineage encompassing all Atlantic samples. In contrast, two Colombian reference mitogenomes formed a separate and highly divergent lineage relative to the Atlantic group, despite showing moderate divergence between them. Comparative mitogenomic analyses revealed conserved genome organization, nucleotide composition bias, codon usage, and transfer RNA (tRNA) structures. All PCGs evolved under strong purifying selection, with Ka/Ks ratios well below unity. CONCLUSIONS: These results support mitochondrial genetic continuity across the Western Atlantic Narcine populations and do not provide mitochondrial evidence for multiple evolutionary lineages within the Western Atlantic. The marked mitochondrial divergence of Colombian reference mitogenomes highlights potential issues in sequence attribution and underscores the importance of data curation. Overall, complete mitochondrial genomes provide a robust framework for species delimitation and future integrative taxonomic assessments within Narcine.

Animals

One-time duplication and ongoing loss of mitochondrial tRNA genes in Cryptocercus cockroaches.

Mitochondrial genome is a popular marker in phylogenetics and species diversity estimations. Mitogenome is relatively compact and conserved, while gene rearrangements were found in some species across various organisms. Models to explain the origin and evolution of gene rearrangement have been proposed but seldom demonstrated; empirical evidence from closely related species is particularly scarce. Here, through an intensive case study of the cockroach genus Cryptocercus Scudder, 1862, we elucidate the evolution of mitochondrial gene order. This study utilized 51 new samples and re-assembled raw reads of 26 published samples. A diversity of rearrangement patterns is recovered, especially in the tRNA gene cluster between ND3 and ND5, which is effectively explained by the duplication - random loss model. Specifically, the entire tRNA gene cluster was duplicated; this duplication is potentially facilitated by chance binding between the 3' end of ND5 gene and the ND3-trnA region during DNA replication. Furthermore, we reveal that one of the gene copies degenerated stochastically across lineages, directly contributing to the observed diversity in gene arrangement. Gene rearrangement patterns are apomorphies for certain clades, providing additional evidence for the inferred phylogeny and serving as potential indicators of species. This study underscores the importance of intensive sampling and rigorous data curation for deciphering the evolutionary mechanisms.

Duplication–random loss model

Meta-QTL Analysis Reveals Consensus Genomic Regions and Candidate Genes for Resistance to Sudden Death Syndrome in Soybean.

Sudden death syndrome (SDS), caused by Fusarium virguliforme, is one of the most economically important diseases limiting soybean production worldwide. Although numerous quantitative trait loci (QTL) associated with SDS resistance have been reported, inconsistencies among mapping populations, marker systems, and experimental conditions have hindered the identification of robust resistance loci for soybean improvement. In this study, a comprehensive meta-analysis was conducted to integrate published QTL and identify stable consensus genomic regions associated with SDS resistance. After a systematic literature survey and data curation, 153 QTL derived from 14 linkage-mapping studies were analyzed using a custom R-based workflow, resulting in the identification of 23 consensus meta-QTL (MQTL) distributed across 17 chromosomes. Several MQTL, particularly those located on chromosomes 6, 8, 18, and 20, were supported by multiple independent studies and represented major genomic hotspots for SDS resistance. Physical localization and functional annotation of these MQTL identified 217 candidate genes, including genes predicted to be involved in plant defense, signal transduction, transcriptional regulation, and secondary metabolism. Gene Ontology enrichment analysis identified response to salicylic acid as the only biological process that remained significant after FDR correction, whereas Kyoto Encyclopedia of Genes and Genomes pathway analysis did not identify significantly enriched pathways. Independent support using five published genome-wide association studies further supported several MQTL, especially those on chromosomes 6, 18, and 20, thereby increasing confidence in these genomic regions. The identified MQTL and prioritized candidate genes provide potential genomic resources for future marker development, improvement applications, and functional validation aimed at improving soybean resistance to SDS.

Fusarium virguliforme

Sharing and community curation of mass spectrometry data with Global Natural Products Social Molecular Networking.

The potential of the diverse chemistries present in natural products (NP) for biotechnology and medicine remains untapped because NP databases are not searchable with raw data and the NP community has no way to share data other than in published papers. Although mass spectrometry (MS) techniques are well-suited to high-throughput characterization of NP, there is a pressing need for an infrastructure to enable sharing and curation of data. We present Global Natural Products Social Molecular Networking (GNPS; http://gnps.ucsd.edu), an open-access knowledge base for community-wide organization and sharing of raw, processed or identified tandem mass (MS/MS) spectrometry data. In GNPS, crowdsourced curation of freely available community-wide reference MS libraries will underpin improved annotations. Data-driven social-networking should facilitate identification of spectra and foster collaborations. We also introduce the concept of 'living data' through continuous reanalysis of deposited data.

Biological Products

usiGrabber: automating the curation of proteomics spectra data at scale, making large datasets ready for use in machine learning systems.

MOTIVATION: An unprecedented amount of mass spectrometry-based proteomics data is publicly available through repositories such as the PRoteomics IDEntifications Database (PRIDE), and the field is increasingly leveraging machine-learning approaches. However, the available data is not ready to be reused in a scalable way beyond the original acquisition purpose. Existing machine learning models commonly rely on a few manually curated datasets that require deep domain expertise and tedious technical work to construct. Importantly, these datasets have not been updated in recent years, so that newly published data remains inaccessible. We present usiGrabber, a scalable framework for assembling large proteomic datasets. usiGrabber is designed around portability and extensibility. It extracts spectra identification data from mzIdentML files, stores additional project-level metadata retrieved through the PRIDE API, indexes raw spectra using Universal Spectrum Identifiers (USIs), and offers download utilities to retrieve spectra data at scale. RESULTS: Within 49 h, we parsed over 800 million peptide spectrum matches and corresponding USIs from over 1200 projects. As a proof of concept, we used usiGrabber to construct a phosphorylation-specific training dataset of nearly 11 million spectra in under 2 days and used it to retrain a binary phosphorylation classifier based on the AHLF model architecture. With a balanced accuracy of 0.78, our model achieves comparable performance to the original model on an independent test set, showing that automated data extraction is an alternative to manual curation of static datasets. AVAILABILITY AND IMPLEMENTATION: All code is available at https://github.com/usiGrabber/usiGrabber; the data are available at https://zenodo.org/records/18853258.

Machine Learning

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository.

Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata. Here, we introduce scBaseCount, a database that leverages an AI agent to automate discovery and metadata extraction and standardize data processing. Built by mining all 10x Genomics datasets, scBaseCount is the largest public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues. It offers an unbiased view of the data landscape within the SRA and enables the training of more performant computational models through access to broader phenotypic diversity. Uniform processing enables measurement of both intronic and exonic reads and non-coding gene expression and improves alignment across experiments. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to autonomously curate biological data repositories.

Single-Cell Analysis

CLASPP: A unified model for predicting post-translational modifications.

Post-Translational Modifications (PTMs) are a fundamental mechanism for regulating cellular pathways and increasing the functional diversity of the proteome. Accurately predicting the PTM types that are likely to occur at a given site in the primary sequence is a key challenge in functional proteomics. Existing PTM prediction models predominantly focus on either single PTM types or employ ensemble methods that combine multiple models to predict different PTM types. This fragmentation is largely driven by the vast imbalance in data availability across PTM types, making it difficult to predict multiple PTM types with a single model. To address this limitation, we present the Contrastively Learned Attention-based Stratified PTM Predictor (CLASPP), a unified PTM prediction model. CLASPP addresses imbalance challenges by leveraging unsupervised clustering-based undersampling and a novel contrastive learning framework tailored to PTM data. Additionally, our hierarchical data organization and curation are shown to improve CLASPP's performance by balancing the representation of individual PTM types and provides a standardized dataset to train and validate future model designs. Drawing inspiration from advancements in image and natural language processing, the CLASPP model employs a multi-stage training strategy and a high-quality, curated training dataset to improve PTM prediction performance. To uncover what is learned during the contrastive learning stage, the CLASPP model is shown to distinguish known protein kinase substrate specificity profiles as a form of explainability. Finally, we evaluate the application of CLASPP in predicting PTMs in different model organisms and experimentally validated ubiquitination sites in the understudied DCLK3 kinase. Overall, CLASPP represents a unified model for PTM prediction that addresses key bottlenecks in data imbalance and offers new strategies for biological data curation, thereby improving PTM-type prediction performance across diverse organisms.

Protein Processing, Post-Translational

High-accuracy SNV calling for bacterial isolates using deep learning with AccuSNV.

Accurate detection of mutations within bacterial species is critical for fundamental studies of microbial evolution, reconstruction of transmission events, and identification of antimicrobial resistance mutations. Although many tools have been developed to identify single-nucleotide variants (SNVs) from whole-genome sequencing, they often suffer from high false-positive rates owing to the complexity of bacterial genomes and the need for different filtering cutoffs across sample types and sequencing depths. As data sets increase in size, the manual filtering required for high accuracy presents a significant obstacle. Here, we present AccuSNV, a novel deep learning-based tool for high-precision and automated bacterial SNV calling. Unlike traditional methods that process one sample at a time, AccuSNV leverages a convolutional neural network (CNN) that integrates alignment information across multiple samples, enhancing precision through learned across-sample patterns. We evaluate AccuSNV against seven popular SNV-calling tools using simulated data from six bacterial species with varied sequencing depths, numbers of isolates, mutations, and divergence levels. To further validate its real-world utility, we test AccuSNV on multiple curated bacterial data sets containing reported SNVs. In both simulated and real-world scenarios, AccuSNV consistently achieves the best performance. Moreover, AccuSNV provides comprehensive user-friendly downstream analysis modules and outputs, including mutation annotation information, phylogenetic inference, d N/d S calculations, and optional manual filtering. Together with the automated deep learning-based calling, these features make AccuSNV broadly accessible to users with different levels of computational expertise.

Deep Learning

PMkbase (version 1.0): an interactive web-based tool for tracking bacterial metabolic traits using phenotype microarrays made interoperable with sequence information and visualizing/processing PM data.

Bacteria showcase remarkable metabolic diversity and traits, even among strains of the same species. In recent years, a large number of bacterial genomes have been sequenced, leading to the elucidation and documentation of genomic differences and commonalities across and within species. Genome-scale metabolic reconstructions, which are often defined and curated using data from phenotype microarrays, elucidate the differences in metabolic traits resulting from genomic diversity. These microarrays measure cellular respiration on a variety of carbon, nitrogen, phosphorus, and sulfur sources and various stressors and inhibitors over a period of time to determine the metabolic activity of a given strain. Despite their popularity in measuring bacterial metabolic activity and traits, no public databases that allow researchers to warehouse, access, and analyze this information currently exist. Additionally, there are no publicly available tools that allow researchers to view the variance of these metabolic traits across bacterial strains. To address this need, we present Phenotype Microarray Knowledgebase (PMkbase [version 1.0], https://pmkbase.com/), an interactive database that acts as a repository of phenotype microarray (PM) data with integrated sequence information. Binarized activity calls, along with associated kinetic parameters, are made for all metabolic substrates and inhibitors. Users can upload their own data for analysis and visualization and to perform quality checks on their experiments. PMkbase will address an unmet need to track and view bacterial metabolic traits and provide researchers with valuable information to develop metabolic models, enrich pangenomic analyses, and design new experiments.IMPORTANCEBacterial species can be differentiated by their metabolic profiles or the type of nutrients they consume. Interestingly, strains within the same species also display differences in nutrient consumption. Phenotype microarrays are a high-throughput, widely used technology to measure which substrates can be metabolized by various microbial strains and the extent to which inhibitors can affect it. Despite their widespread use, public databases to parse and access this data type at scale do not exist. PMkbase, which contains 9,024 data points for nitrogen substrate utilization, 41,664 data points for carbon substrate utilization, 8,448 data points for phosphorus/sulfur substrate utilization, and 27,264 data points on various antibiotics across three species (Escherichia coli, Pseudomonas putida, and Staphylococcus aureus), has been developed to allow researchers to freely access PM data, along with enriching the data with sequence information.

Bacteria

ClinGen recuration of hearing loss-associated genes demonstrates significant changes in gene-disease validity over time.

PURPOSE: The Clinical Genome Resource (ClinGen) Hearing Loss Gene Curation Expert Panel was assembled in 2016 and has since curated 174 gene-disease relationships (GDRs) using ClinGen's semiquantitative framework. ClinGen mandates the timely recuration of all GDRs classified as Disputed, Limited, Moderate, and Strong every 2 to 3 years. METHODS: Thirty-five GDRs met the criteria for recuration within 2 years of original curation. Previous evidence was reevaluated using the latest curation guidelines, and a comprehensive literature review was performed to obtain new evidence. Recurations were approved by the Gene Curation Expert Panel and published on the ClinGen website (www.clinicalgenome.org). RESULTS: Eight of 35 GDRs (22%) changed their classification. Two Moderate and 5 Strong GDRs were upgraded to Definitive because of new case evidence. One Strong was subsumed under another Definitive GDR after evaluation of the lumping/splitting of disease entities. Twenty-seven of 35 patients remained unchanged, with little to no new evidence reported. CONCLUSION: Genes classified as Moderate and Strong were likely to build evidence and change their classification over time, whereas Limited were unlikely to gain evidence. These findings highlight the critical role of recuration in ensuring that genetic tests and research studies incorporate the most recent evidence into their efforts.

Humans

Retinal Proteome Profiling of Inherited Retinal Degeneration Across Three Different Mouse Models Suggests Common Drug Targets in Retinitis Pigmentosa.

Inherited retinal degenerations (IRDs) are a leading cause of blindness among the population of young people in the developed world. Approximately half of IRDs initially manifest as gradual loss of night vision and visual fields, characteristic of retinitis pigmentosa (RP). Due to challenges in genetic testing, and the large heterogeneity of mutations underlying RP, targeted gene therapies are an impractical largescale solution in the foreseeable future. For this reason, identifying key pathophysiological pathways in IRDs that could be targets for mutation-agnostic and disease-modifying therapies (DMTs) is warranted. In this study, we investigated the retinal proteome of three distinct IRD mouse models, in comparison to sex- and age-matched wild-type mice. Specifically, we used the Pde6βRd10 (rd10) and RhoP23H/WT (P23H) mouse models of autosomal recessive and autosomal dominant RP, respectively, as well as the Rpe65-/- mouse model of Leber's congenital amaurosis type 2 (LCA2). The mice were housed at two distinct institutions and analyzed using LC-MS in three separate facilities/instruments following data-dependent and data-independent acquisition modes. This cross-institutional and multi-methodological approach signifies the reliability and reproducibility of the results. The large-scale profiling of the retinal proteome, coupled with in vivo electroretinography recordings, provided us with a reliable basis for comparing the disease phenotypes and severity. Despite evident inflammation, cellular stress, and downscaled phototransduction observed consistently across all three models, the underlying pathologies of RP and LCA2 displayed many differences, sharing only four general KEGG pathways. The opposite is true for the two RP models in which we identify remarkable convergence in proteomic phenotype even though the mechanism of primary rod death in rd10 and P23H mice is different. Our data highlights the cAMP and cGMP second-messenger signaling pathways as potential targets for therapeutic intervention. The proteomic data is curated and made publicly available, facilitating the discovery of universal therapeutic targets for RP.

Animals

PAHG: the database of human multi-gene families.

BACKGROUND: In the early vertebrate history, gene duplications, including single-gene, segmental-gene (SSD), and whole-genome duplication (WGD), formed multigene families. Despite efforts to classify metazoan multigene families hierarchically for evolutionary insight, a gap exists in accessible, curated resources for human/vertebrate multigene families. RESULTS: Addressing this, we present the Phylogenomic Analysis of Human Genome (PAHG) database. It focuses on curated multigene families in the human genome, particularly within four paralogons: HOX-bearing (Hsa:2/7/12/17), FGFR-bearing (Hsa:4/5/8/10), MHC-bearing (Hsa:1/6/9/19), and chromosomes 1/2/8/20. CONCLUSION: The current PAHG version details the phylogenetic history of 221 human multigene families (1247 gene members) with 15,231 protein sequences from diverse metazoans. It provides insights into gene duplication timings, co-duplication events, and their relationships with human genome syntenic organization. The PAHG database addresses the lack of accessible resources, offering valuable information on human/vertebrate multigene family evolution. Access the PAHG database at: https://www.pahgncb.com/ and http://pahg.qau.edu.pk/ . This resource enriches our understanding of vertebrate genetic evolution.

Humans

Unveiling microbial risks in Chinese household dust: a comprehensive analysis from absolute abundance to virulence unit.

BACKGROUND: People spend the majority of their lives indoors, yet the risk and virulence potential of household microbiota remain largely unexplored, particularly in developing countries. RESULTS: Here, we conducted a nationwide survey on both dust samples and health information across 118 Chinese households. The microbiota composition and its functional units were analyzed using absolute 16S rRNA/ITS sequencing, metagenomics, and metaproteomics. Cross-domain network analysis of the core microbial communities revealed robust co-occurrence patterns in household dust. The mean absolute abundance of potentially pathogenic bacteria and fungi in households was 2.39 × 105 and 2.83 × 106 DNA copies/g dust. The potentially pathogenic community was primarily influenced by latitude, relative humidity, and average temperature. Although total absolute abundance was substantially lower in urban areas, the relative abundance of potentially pathogenic bacteria was markedly higher compared to rural environments. While urban-rural differences existed, the underlying statistical drivers were the environmental variables. The absolute abundance of potential pathogens was significantly associated with the prevalence of rhinitis, wheeze, and dermatitis in 266 participants. Children were identified as the highest-risk group from inhalation exposure of average daily dose. A total of 170 bacterial, 223 fungal virulence factors (VFs), and 370 antibiotic resistance genes (ARGs) were detected in dust and dust extracellular vesicle (EV)-associated DNA. EV-associated cargoes contributed 47.13% to the bacterial VF profiles, 11.90% to fungal VF profiles, and 44.45% to ARG profiles. Metaproteomic analysis confirmed the presence of VF profiles in dust EVs, which was further verified by curated proteomics data from 35 household pathogens. CONCLUSIONS: This study provides a comprehensive, quantitative framework linking indoor microbial exposure to health risks, highlighting EVs as a non-negligible, novel, extracellular mechanistic pathway for health impact in household environments. Video Abstract.

Child

Large-scale functional annotation establishes a reference framework for human LRRK2 variants.

Pathogenic variants in leucine-rich repeat kinase 2 (LRRK2)1are among the most frequent monogenic causes of Parkinson's disease (PD)2 and act through a gain-of-function mechanism of increased kinase activity. LRRK2-targeted therapies are in clinical development, but interpretation of the rapidly expanding catalogue of rare LRRK2 variants remains a barrier to translation. Here, we present functionally annotated data on >350 LRRK2 coding variants using a standardized cellular assay with Rab10 phosphorylation as a readout of kinase activity and integrated these data with curated genetic and clinical annotations from the Movement Disorders Society Genetic Mutation Database (MDSGene). Variants differed in activation magnitude, ranging from modest increases (e.g., p.G2019S) to strongly activating substitutions such as p.Y1699C or p.L1795F. Activating variants occurred across the full length of LRRK2, although the largest effects clustered within the ROC-COR regulatory hub, where structural analysis identified subdomains forming an allosteric scaffold controlling kinase output. All known/established pathogenic variants showed increased activity, whereas benign and likely benign variants remained within the wild-type range. Functional effect sizes correlated with pathway activation in patient-derived immune cells, altogether providing a framework for ACMG-based variant interpretation in which kinase activation can support PS3 functional evidence for reclassification of variants.

Protein phosphorylation

A decentralized future for the open-science databases.

The continuous and reliable open access to curated biological data repositories is indispensable for accelerating rigorous scientific inquiry and fostering reproducible research outcomes. However, the current paradigm, which relies heavily on centralized infrastructure for the storage and distribution of foundational biomedical datasets, inherently introduces significant vulnerabilities. This centralized model is susceptible to single points of failure, including cyberattacks, technical malfunctions, natural disasters, and even political or funding uncertainties. Such disruptions can lead to widespread data unavailability, data loss, integrity compromises, and substantial delays in critical research, ultimately impeding scientific progress. The downstream effect of such interruptions can be the widespread paralysis of diverse research activities, including computational, clinical, molecular, and climate studies. This scenario vividly illustrates the inherent dangers of consolidating essential scientific resources within a single geopolitical or institutional locus. As data generation is accelerating and the global landscape continues to fluctuate, the sustainability of centralized models must be critically re-evaluated. A shift toward federated and decentralized architectures may offer a robust and forward-looking approach to enhancing the resilience of scientific data infrastructures by reducing exposure to governance instability, infrastructural fragility, and funding volatility, while also promoting equity and global accessibility. Inspired by established models such as ELIXIR's federated infrastructure and the policy and funding frameworks developed by CODATA and the Global Biodata Coalition (GBC), emerging Decentralized Science (DeSci) initiatives can contribute to building more resilient, fair, and incentive-aligned data ecosystems. The future of open science depends on integrating these complementary approaches to establish a globally distributed, economically sustainable, and institutionally robust infrastructure that safeguards scientific data as a public good, further ensuring continued accessibility, interoperability, and preservation for generations to come. Here, we examine the structural limitations of centralized repositories, evaluate federated and decentralized models, and propose a hybrid framework for resilient, fair, and sustainable scientific data stewardship.

data accessibility