PubMed HealthSearch

SEARCH · PubMed Health

Results for “open resource”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

MetaflowX: a scalable and resource-efficient workflow for multi-strategy metagenomic analysis.

Microbiomes play crucial roles in diverse ecosystems, spanning environmental, agricultural, and human health domains. However, in-depth metagenomic data analysis presents significant technical and resource challenges, particularly at scale. Existing computational pipelines are typically limited to either reference-based or reference-free approaches and exhibit inefficiencies in process large datasets. Here, we introduce MetaflowX (https://github.com/01life/MetaflowX), an open-resource workflow integrating both analytical paradigms for enhanced metagenomic investigations. This modular framework encompasses short-read quality control, rapid microbial profiling, hybrid contig assembly and binning, high-quality metagenome-assembled genome (MAG) identification, as well as bin refinement and reassembly. Benchmarking tests showed that MetaflowX completed full metagenomic analyses up to 14-fold faster and with 38% less disk usage than existing workflows. It also recovered the highest number of high-quality and taxonomically diverse MAGs. A dedicated reassembly module further improved MAG quality, increasing completeness by 5.6% and reducing contamination by 53% on average. Functional annotation modules enable detection of key features, including virulence and antibiotic resistance genes. Designed for extensibility, MetaflowX provides an efficient solution addressing current and emerging demands in large-scale metagenomic research.

Metagenomics

Pan-cancer multi-omics machine learning defines a lactylation-associated immune-excluded tumor state with proteomic and experimental corroboration.

BACKGROUND: Histone lactylation links lactate metabolism to chromatin regulation, but whether lactylation-program-associated transcriptional patterns delineate recurrent pan-cancer tumor states remains unclear. METHODS: We integrated mRNA, lncRNA, and miRNA profiles from 9712 TCGA tumors across 33 cancer types with GTEx references, six GEO cohorts, IMvigor210, and an institutional clear-cell renal cell carcinoma (ccRCC) cohort used for exploratory DIA-NN proteomic corroboration. Random-effects co-expression meta-analysis, multi-omics consensus clustering, regulon inference, immune deconvolution, TIDE, oncoPredict, and SHAP-based machine learning were applied. hsa-miR-431-5p was functionally evaluated as a proof-of-concept CS2-associated miRNA in bladder cancer models. RESULTS: LacCoEx-Atlas comprised 398,491 lactylation-related co-expression pairs across 24,667 RNA features under a random-effects framework (median I² = 88.6%). Consensus clustering identified two subtypes: CS2 showed glycolytic-mesenchymal-immune-excluded features, M2 macrophage enrichment, CD8⁺ T-cell depletion, elevated HDAC4/NSD3/KDM6B activity, and worse survival, whereas CS1 showed oxidative, sirtuin-active programs. CS2 had fewer predicted ICI responders (18.3% vs. 52.0%) and a lower observed ORR in IMvigor210 (15.3% vs. 24.0%). oncoPredict identified NU7441 as a hypothesis-generating CS2-associated sensitivity signal (Hedges' g = 1.17). DIA-NN proteomics in 50 ccRCC specimens provided exploratory support for CS2-associated hypoxia, ECM degradation, and metastasis programs. The 10-feature mRNA LARItools model achieved an apparent AUC of 0.9413, while a separate multi-omics model achieved 0.971; neither was independently validated. LARItools reproduced prognostic separation across six GEO cohorts. miR-431-5p promoted malignant phenotypes and EMT in bladder cancer cells, with concordant CMU4h expression findings. CONCLUSIONS: Lactylation-program-associated transcriptional patterns delineate a recurrent immune-excluded pan-cancer tumor state associated with adverse prognosis, reduced predicted immunotherapy responsiveness, exploratory single-cancer protein-level support, and testable DNA damage response-targeting hypotheses. LacCoEx-Atlas and LARItools provide open resources for lactylation-program-associated tumor-state stratification and future translational research.

Humans

Whole-genome sequencing of 490,640 UK Biobank participants.

Whole-genome sequencing provides an unbiased and complete view of the human genome and enables the discovery of genetic variation without the technical limitations of other genotyping technologies. Here we report on whole-genome sequencing of 490,640 UK Biobank participants, building on previous genotyping effort1. This advance deepens our understanding of how genetics associates with disease biology and further enhances the value of this open resource for the study of human biology and health. Coupling this dataset with rich phenotypic data, we surveyed within- and cross-ancestry genomic associations and identified novel genetic and clinical insights. Although most associations with disease traits were primarily observed in individuals of European ancestries, strong or novel signals were also identified in individuals of African and Asian ancestries. With the improved ability to accurately genotype structural variants and exonic variation in both coding and UTR sequences, we strengthened and revealed novel insights relative to whole-exome sequencing2,3 analyses. This dataset, representing a large collection of whole-genome sequencing data that is available to the UK Biobank research community, will enable advances of our understanding of the human genome, facilitate the discovery of diagnostics and therapeutics with higher efficacy and improved safety profile, and enable precision medicine strategies with the potential to improve global health.

Humans

A custom library construction method for super-resolution ribosome profiling in Arabidopsis.

BACKGROUND: Ribosome profiling, also known as Ribo-seq, is a powerful technique to study genome-wide mRNA translation. It reveals the precise positions and quantification of ribosomes on mRNAs through deep sequencing of ribosome footprints. We previously optimized the resolution of this technique in plants. However, several key reagents in our original method have been discontinued, and thus, there is an urgent need to establish an alternative protocol. RESULTS: Here we describe a step-by-step protocol that combines our optimized ribosome footprinting in plants with available custom library construction methods established in yeast and bacteria. We tested this protocol in 7-day-old Arabidopsis seedlings and evaluated the quality of the sequencing data regarding ribosome footprint length, mapped genomic features, and the periodic properties corresponding to actively translating ribosomes through open resource bioinformatic tools. We successfully generated high-quality Ribo-seq data comparable with our original method. CONCLUSIONS: We established a custom library construction method for super-resolution Ribo-seq in Arabidopsis. The experimental protocol and bioinformatic pipeline should be readily applicable to other plant tissues and species.

3-nt periodicity

The ASH HematOmics Program supports integrative analysis of genomic and clinical data in hematologic diseases.

The increasing availability of genomic and transcriptomic sequencing has uncovered diverse genomic alterations and distinct gene expression profiles driving hematologic diseases, yet a data integration and sharing platform dedicated to hematology remains lacking. We developed the American Society of Hematology (ASH) HematOmics Program (ASHOP; ashop.hematology.org), a resource for exploring somatic alterations and gene fusions, transcriptomic results, and clinical data from 5960 patients spanning B-cell precursor and T-cell acute lymphoblastic leukemia, acute myeloid leukemia, myelodysplastic syndromes, and chronic lymphocytic leukemia. Users can explore genomic alteration landscapes and comutation patterns via lollipop and matrix plots and analyze significantly altered genes in user-defined subcohorts. Transcriptomes can be explored through interactive uniform manifold approximation and projections, clustering, differential expression, and pathway enrichment. Genomic, transcriptomic features, and clinical outcomes can be correlated in a user-driven manner or combined to precisely define study cohorts. We illustrate the following 4 use cases of ASHOP: (1) stratification of DUX4-rearranged B-cell leukemias into Early/Multipotent and Committed subgroups with distinct outcomes, (2) characterization of HOXA/HOXB expression patterns in acute myeloid leukemias, (3) correlating mutational burden with mismatch repair deficiency and mutational signatures, and (4) investigation of TP53 alteration landscape. ASHOP is an open-access resource to inform genomic and transcriptomic data interpretation for hematologic malignancies and will expand to support additional diseases and data modalities from the ASH community.

Humans

Associations on the Fly, a new feature aiming to facilitate exploration of the Open Targets Platform evidence.

MOTIVATION: The Open Targets Platform (https://platform.opentargets.org) is a unique, comprehensive, open-source resource supporting systematic identification and prioritisation of targets for drug discovery. The Platform combines, harmonizes and integrates data from >20 diverse sources to provide target-disease associations, covering evidence derived from genetic associations, somatic mutations, known drugs, differential expression, animal models, pathways and systems biology. An in-house target identification scoring framework weighs the evidence from each data source and type, contributing to an overall score for each of the 7.8M target-disease associations. However, the old infrastructure did not allow user-led dynamic adjustments in the contribution of different evidence types for target prioritisation, a limitation frequently raised by our user community. Furthermore, the previous Platform user interface did not support navigation and exploration of the underlying target-disease evidence on the same page, occasionally making the user journey counterintuitive. RESULTS: Here, we describe 'Associations on the Fly' (AOTF), a new Platform feature-developed with a user-centred vision-that enables the user to formulate more flexible therapeutic hypotheses through dynamic adjustment of the weight of contributing evidence from each source, altering the prioritisation of targets. AVAILABILITY AND IMPLEMENTATION: The codebases that power the Platform-including our pipelines, GraphQL API, and React UI-are all open source and licensed under the APACHE LICENSE, VERSION 2.0. You can find all of our code repositories on GitHub at https://github.com/opentargets and on Zenodo at https://zenodo.org/records/14392214. This tool was implemented using React v18 and its code is accessible here: (https://github.com/opentargets/ot-ui-apps). The tools are accessible through the Open Targets Platform web interface (https://platform.opentargets.org/) and GraphQL API (https://platform-docs.opentargets.org/data-access/graphql-api). Data is available for download here: (https://platform.opentargets.org/downloads) and from the EMBL-EBI FTP: (https://ftp.ebi.ac.uk/pub/databases/opentargets/platform/).

Software

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals

Development and extensive sequencing of a broadly-consented Genome in a Bottle matched tumor-normal pair.

The Genome in a Bottle Consortium (GIAB), hosted by the National Institute of Standards and Technology (NIST), is developing new matched tumor-normal samples, the first explicitly consented for public dissemination of genomic data and cell lines. Here, we describe a comprehensive genomic dataset from the first individual, HG008, including DNA from an adherent, epithelial-like pancreatic ductal adenocarcinoma (PDAC) tumor cell line and matched normal cells from duodenal and pancreatic tissues. Data for the tumor-normal matched samples comes from seventeen distinct state-of-the-art whole genome measurement technologies, including high depth short and long-read bulk whole genome sequencing (WGS), single cell WGS, Hi-C, and karyotyping. These data will be used by the GIAB Consortium to develop matched tumor-normal benchmarks for somatic variant detection. We expect these data to facilitate innovation for whole genome measurement technologies, de novo assembly of tumor and normal genomes, and bioinformatic tools to identify small and structural somatic variants. This first-of-its-kind broadly consented open-access resource will facilitate further understanding of sequencing methods used for cancer biology.

Humans

An educator framework for organizing Wikipedia editathons for computational biology.

MOTIVATION: Wikipedia is a vital open educational resource in computational biology; however, a significant knowledge gap exists between English and non-English Wikipedias. Reducing this knowledge gap via intensive editing events, or "editathons," would be beneficial in reducing language barriers that disadvantage learners whose native language is not English. Results: We present a framework to guide educators in organizing editathons for learners to improve and create relevant Wikipedia articles. As a case study, we present the results of an editathon held at the 2024 ISCB Latin America conference, in which ten new articles were created for the Spanish-language edition of Wikipedia. We also present a web tool, "compbio-on-wiki," which identifies relevant English Wikipedia articles missing in other languages. We demonstrate the value of editathons to expand the accessibility and visibility of computational biology content in multiple languages. AVAILABILITY AND IMPLEMENTATION: Source code for the compbio-on-wiki Toolforge site is available at: https://github.com/lubianat/compbio-on-wiki.

Computational Biology

Development and extensive sequencing of a broadly-consented Genome in a Bottle matched tumor-normal pair.

The Genome in a Bottle Consortium (GIAB), hosted by the National Institute of Standards and Technology (NIST), is developing new matched tumor-normal samples, the first to be explicitly consented for public dissemination of genomic data and cell lines. Here, we describe a comprehensive genomic dataset from the first individual, HG008, including DNA from an adherent, epithelial-like pancreatic ductal adenocarcinoma (PDAC) tumor cell line and matched normal cells from duodenal and pancreatic tissues. Data for the tumor-normal matched samples comes from seventeen distinct state-of-the-art whole genome measurement technologies, including high depth short and long-read bulk whole genome sequencing (WGS), single cell WGS, and Hi-C, and karyotyping. In future publications, these data will be used by the GIAB Consortium to develop matched tumor-normal benchmarks for somatic variant detection. We expect these data to facilitate innovation for whole genome measurement technologies, de novo assembly of tumor and normal genomes, and bioinformatic tools to identify small and structural somatic mutations. This first-of-its-kind broadly consented open-access resource will facilitate further understanding of sequencing methods used for cancer biology.

Journal Article

Exploring endothelial cell environments across organs in spatially resolved omics data.

Endothelial cells are ubiquitously present in the human body and line the luminal surface of blood and lymphatic vessels. The oxygen-dependence of cells impacts their proximity to blood vessels, and consequently, to endothelial cells depending on their functional properties and priorities. This paper presents cell-to-nearest-endothelial-cell distance distributions for various cell types using 399 spatially resolved omics datasets from 14 studies comprising 12 tissue types with a total of 47,349,496 cells. Additionally, we developed an open-source web-based interactive tool, Cell Distance Explorer, that allows researchers to interactively visualize cell graphs and linkages in 2D and 3D datasets. Finally, we present a hierarchical neighborhood analysis focused on the endothelial cell neighborhoods in small and large intestine datasets. This paper provides an open-access resource (datasets, tools, and analyses) to characterize and compare cell distances and cell neighborhoods in spatially resolved omics data.

Journal Article

Comparative transcriptomics uncovers poplar and fungal genetic determinants of ectomycorrhizal compatibility.

Ectomycorrhizal symbiosis supports tree growth and is crucial for nutrient cycling and temperate and boreal ecosystems functioning. The establishment of functional ectomycorrhiza (ECM) first requires the association of compatible partners. However, host and fungal genetic determinants governing mycorrhizal compatibility are unknown. To identify such factors in poplar and its fungal associates, we mined existing and de novo tree and fungal transcriptional datasets. We identified co-expressed genes enabling ECM symbiosis at early and mature stages of the interaction. These sets of genes can be divided into general fungal-sensing and ECM-specific components. We highlight the importance of fungal modulation of plant JA-related defenses and the regulation of secretory pathways for ECM compatibility, including upregulation of key fungal small secreted proteins, the downregulation of plant secreted peroxidases, and the downregulation of plant cell wall remodeling proteins concomitantly with the upregulation of fungal glycosyl hydrolases acting on pectin. Not only gene regulation, but also its temporal scale and dynamics seem to play a crucial role for mycorrhizal compatibility. The expression profile of the host Common Symbiosis Pathway and nutrient transporters was also studied, revealing constitutive levels of expression and moderate upregulation in compatible ECM interactions. Overall, these results underscore the importance of novel biological functions during the establishment of ECM symbiosis, help us gain insights into the molecular events determining mycorrhiza compatibility, and serve as a data-rich transcriptomic resource to open new research questions in the field.

Mycorrhizae

Genome wide association study of rice agronomical traits and seed ionome with the NARO Open Rice Collection.

To meet the nutritional needs of the rising human population, genetic variants are necessary for the breeding of new cultivars. Rice (Oryza sativa L.) is a staple food for over half of the world's population. Here, we developed a new rice genetic resource, the NARO Open Rice Collection (NRC) with high-resolution genome data. NRC consists of 623 accessions, and approximately 200 accessions are categorized into three major subgroups, categorized as Indica, Japonica, and Aus. In this study, we performed genome-wide association studies (GWAS) for rice heading date, seed shape, and seed ionome using the NRC. Well-known genes related to heading date and seed shape were detected by GWAS using the NRC accessions. Therefore, we concluded that our new rice collection is suitable for GWAS. In addition, GWAS with each subgroup was advantageous for the detection of particular genes. Finally, we performed GWAS for seed ionome with the aim of improving the nutritional properties of rice, as essential minerals for humans, such as iron (Fe) and zinc (Zn), are not sufficient in rice seeds. Our study revealed that OsATL31, a likely ubiquitin E3 ligase, was involved in the control of Fe and Zn contents in seeds.

Oryza

Probiogenomic analysis of functional potential and safety of L. plantarum 8p-a3 and DMC-S1 strains: in silico vs in vitro and in vivo data.

The molecular basis of the beneficial effects and the causes of the negative effects of probiotics are not entirely clear. Clarifying these issues is important for understanding the biology and assessing the safety of the microbes. Omics technologies have opened up new resources for obtaining relevant knowledge. Here, for the first time, we present the results of a comparative analysis of the functional potential and safety of two L. plantarum strains: the approved probiotic 8p-a3 and the Drosophila intestinal resident, which exhibit opposite effects on D. melanogaster as the model host organism. Through genomic analysis, extracellular vesicle studies, and in vitro and in vivo assays, we have identified the common and specific characteristics of the strains. The strains proved to be similar in a set of genes that determine benefits to the host organism, as well as in the presence of some risk factors. Significant differences between the strains are related to genes responsible for adhesion, sialic acid metabolism, mucin degradation, antimicrobial peptides, tannin resistance, and immunomodulation. In silico data correlated with in vitro and in vivo data, with the exception of antimicrobial sensitivity. Pronounced differences between the strains were found in terms of the composition and biological effects of their vesicles. In vivo data on the effects of the strains correlate with the corresponding data of their vesicles in the fruit fly model. The results obtained open up new facets in L. plantarum strains relevant for evaluating the functionality and safety of probiotics.IMPORTANCEUsing a probiogenomic approach, common and specific features regarding functionality and safety were identified in the strains (the approved probiotic strain L. plantarum 8p-a3 and the Drosophila intestinal bacterium L. plantarum DMC-S1), which exhibit opposite effects on the model host organism (D. melanogaster). The genomic analysis was supplemented by the analysis of extracellular vesicles of the strains. Comparative analysis of in silico data in combination with in vitro and in vivo studies was performed, and unexpected capabilities of the strains were discovered. Novel factors, essential for evaluating the safety of probiotics, were identified. New facets in the interplay of probiotic bacterium with host organism have been revealed.

Animals

Multi-ancestry multi-trait analysis reveals shared genetics across major psychiatric disorders and Alzheimer's disease.

The clinical overlap between major psychiatric disorders (MPDs) and Alzheimer's disease (AD) implicates complex shared etiology. Previous studies demonstrated that both diseases are genetically complex and highly heritable, suggesting that more endeavors are necessary to be made from the very bottom to understand their genetic basis. With the advance of post-genomic analysis, multi-ancestry meta-analysis allows the generalizability of the genetic architecture across different populations to uncover ancestry-specific variants, while multi-trait analysis enables the discovery of the co-colocalized risk genomic regions across diseases. Therefore, in this study, we leveraged published GWAS summary statistics from European, East Asian, Hispanic and African American populations to report schizophrenia, major depressive disorders, and Alzheimer's disease risk loci and further fine-mapping to credible sets with >95% PP inclusion of the causal variant. We distilled 2871 potential traits from publicly available and found 134 traits significantly genetically correlated with both MPDs and AD using batch LD score regression. We then prioritized the identified loci from multi-ancestry results for cross-trait colocalization analysis to assess shared genetic etiology and further nominated 2 colocalized loci across both conditions, including rs2532240 and rs6504163. In the end, we finalized our analysis by validation and functional inference of the underlying susceptibility genes as well as putative mechanisms using evidence from multiple resources, including FIVEx, Open Targets, and scQTLbase.

Humans

A guide to selecting high-performing antibodies for TMEM175 (UniProt ID: Q9BSA9) for use in western blot, immunoprecipitation, and immunofluorescence.

TMEM175 is the pore-forming subunit of a lysosomal K+ channel complex that regulates lysosomal pH stability and membrane potential. To further investigate its cellular functions and implications in neurodegenerative diseases, antibody reagents are needed. Here we have characterized six TMEM175 commercial antibodies for western blot, immunoprecipitation, and immunofluorescence using a standardized experimental protocol based on comparing read-outs in knockout cell lines and isogenic parental controls. These studies are part of a larger, collaborative initiative seeking to address antibody reproducibility issues by characterizing commercially available antibodies for human proteins and publishing the results openly as a resource for the scientific community. While use of antibodies and protocols vary between laboratories, we encourage readers to use this report as a guide to select the most appropriate antibodies for their specific needs.

Humans

A guide to selecting high-performing antibodies for DJ-1 ( PARK7) (Q99497) for use in western blot, immunoprecipitation, and immunofluorescence.

DJ-1 is a multifunctional protein that plays a pivotal role in cellular protection against oxidative stress and neurodegeneration. Mutations in the PARK7 gene are associated with early-onset familial Parkinson's disease. Here we have characterized sixteen DJ-1 commercial antibodies for western blot, immunoprecipitation, and immunofluorescence using a standardized experimental protocol based on comparing read-outs in knockout cell lines and isogenic parental controls. These studies are part of a larger, collaborative initiative seeking to address antibody reproducibility issues by characterizing commercially available antibodies for human proteins and publishing the results openly as a resource for the scientific community. While the use of antibodies and protocols vary between laboratories, we encourage readers to use this report as a guide to select the most appropriate antibodies for their specific needs.

Protein Deglycase DJ-1

A guide to selecting high-performing antibodies for Syntenin-1 (O00560) for use in western blot, immunoprecipitation, and immunofluorescence.

Syntenin-1 is the Syndecan-binding protein 1 and a PDZ domain-containing adaptor protein that regulates diverse cellular processes through its interactions with transmembrane receptors, cytoskeletal components, and signaling molecules. Here we have characterized twelve Syntenin-1 commercial antibodies for western blot, immunoprecipitation, and immunofluorescence using a standardized experimental protocol based on comparing read-outs in knockout cell lines and isogenic parental controls. These studies are part of a larger, collaborative initiative seeking to address antibody reproducibility issues by characterizing commercially available antibodies for human proteins and publishing the results openly as a resource for the scientific community. While the use of antibodies and protocols vary between laboratories, we encourage readers to use this report as a guide to select the most appropriate antibodies for their specific needs.

Syntenins