PubMed HealthSearch

SEARCH · PubMed Health

Results for “data commons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Institutional data commons: a federated Data Use Certification-aware architecture for secure and scalable data use in biomedical data ecosystems.

BACKGROUND: Modern biomedical data ecosystems increasingly rely on global cloud platforms to coordinate access to large-scale genomic and clinical datasets. However, operational governance remains largely investigator-centric, shifting the responsibility for complex security, compliance, and infrastructure management to individual laboratories. As data volumes and regulatory requirements expand, this approach fails to scale across the research enterprise. This disjointed approach creates a substantial governance burden and can slow down scientific progress. In centralized cloud environments, investigators face siloed identity management and high costs, leading to inefficient data use and increased risk when integrating local and global datasets. MATERIALS AND METHODS: We examine limitations in the current infrastructure and propose reframing institutional data commons as governance-aware intermediaries to ensure secure, efficient and sustainable use of controlled-access biomedical data. RESULTS: This federated architecture decouples storage from authorization, enabling dynamic access linked to active certifications, whether data are analyzed in situ on global platforms or in local governance-aware institutional access environments. DISCUSSION: Shifting governance from investigators to institutional infrastructure ensures that biomedical research remains both secure and economically sustainable.

biomedical data ecosystems

CrossAttOmics: multiomics data integration with cross-attention.

MOTIVATION: Advances in high throughput technologies enabled large access to various types of omics. Each omics provides a partial view of the underlying biological process. Integrating multiple omics layers would help have a more accurate diagnosis. However, the complexity of omics data requires approaches that can capture complex relationships. One way to accomplish this is by exploiting the known regulatory links between the different omics, which could help in constructing a better multimodal representation. RESULTS: In this article, we propose CrossAttOmics, a new deep-learning architecture based on the cross-attention mechanism for multiomics integration. Each modality is projected in a lower dimensional space with its specific encoder. Interactions between modalities with known regulatory links are computed in the feature representation space with cross-attention. The results of different experiments carried out in this article show that our model can accurately predict the types of cancer by exploiting the interactions between multiple modalities. CrossAttOmics outperforms other methods when there are few paired training examples. Our approach can be combined with attribution methods like LRP to identify which interactions are the most important. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/Sanofi-Public/CrossAttOmics and https://doi.org/10.5281/zenodo.15065928. TCGA data can be downloaded from the Genomic Data Commons Data Portal. CCLE data can be downloaded from the depmap portal.

Humans

Accelerographic, temporal, and distance gait factors in below-knee amputees.

Gait characteristics of 19 patients with a unilateral below-knee amputation were studied. The accelerographic and foot placement method used in this study allowed for simultaneous acquisition of data commonly obtained in the experimental laboratory (acceleration) and data easily gathered in the physical therapy clinic (temporal and distance factors). The following results may be of interest to the clinician: 1) measures of cadence, stride length, and velocity were highly related and the magnitude of these measures was below commonly accepted values for normal; 2) the below-knee amputees spent more time in stance phase on their uninvolved lower extremity than on their involved (prosthetic) extemity; 3) the step length from heel strike of the uninvolved lower extremity to heel strike of the involved (prosthetic) lower extremity was greater and accomplished in less time than the opposite step; and 4) smoothness of the gait pattern and any single temporal and distance factor exhibited low statistical relationships.

Acceleration

Two blind spots in the demographic inference of human origins from genomic data.

Ancient DNA and new inference methods have transformed the study of human origins, but consensus has not followed. Evidence increasingly indicates that hominin populations were pervasively structured and admixed, so complexity rather than simplicity is the appropriate prior. Here I highlight two blind spots that impede resolving that complexity. First, every inference passes through summaries of the data, and those summaries bound what can be recovered. Second, the space of candidate models is vast, yet competing model classes are rarely fit to common data, so a reported best model carries little evidence about untested model classes. This second blind spot reflects practice rather than data. It can be narrowed by testing competing models against withheld summaries and by reporting the models that were tried and rejected rather than only the winner.

Journal Article

ClarID: A Human-Readable and Compact Identifier Specification for Biomedical Metadata Integration.

BACKGROUND: In biomedical research, subjects and biospecimens are commonly tracked using simple IDs or UUIDs, which guarantee uniqueness but convey no embedded semantic information. Contextual metadata (such as tissue type, diagnosis, or assay) is often stored separately, making integration, cohort selection, and downstream analysis cumbersome. While structured barcoding systems exist in large consortia (e.g., TCGA, GTEx) or domain-specific contexts (e.g., SPREC, GOLD), no unified, extensible framework currently spans both subjects and biosamples in a human- and machine-readable way. METHODS: We developed ClarID, a domain-agnostic specification that supports two identifier formats: (i) a human-readable form (e.g., 'CNAG_Test-HomSap-00001-LIV-TUM-RNA-C22.0-TRT-P1W' that encodes key metadata such as project, species, subject_id, tissue, assay, disease, timepoint and duration (from that event); and (ii) a compact version named 'stub' (e.g., 'CT01001LTR0N401T1W') optimized for filenames, pipelines, and labeling.ClarID is implemented through an open-source command-line tool, ClarID-Tools, which processes tabular metadata files (CSV/TSV) and uses a YAML-based codebook to generate, decode, and validate identifiers, as well as to create and read QR codes. The tool supports bulk and single-sample processing and allows easy integration with institutional workflows. RESULTS: To demonstrate ClarID's utility, we applied it to datasets from the Genomic Data Commons (GDC), generating interpretable identifiers for more than 113,000 clinical records (subjects) and 4,255 biospecimen records. All materials, including pre-processing scripts, input and encoded data, are publicly available and fully reproducible via the accompanying GitHub repository and Google Colab. CONCLUSIONS: ClarID fills a critical gap between opaque accession numbers and rich metadata schemas by embedding key context directly into structured identifiers. It enhances traceability, facilitates downstream analysis, and remains adaptable to project-specific needs through a configurable codebook. The accompanying ClarID-Tools software is freely available, together with full documentation and reproducible pipelines, at https://github.com/CNAG-Biomedical-Informatics/clarid-tools.

Biosample identifiers

[Functional stability of the cerebral circulatory system].

Functional stability of the cerebral circulation system seems to be based on the active mechanisms and on those stemming from specifics of the biophysical structure of the system under study. This latter parameter has some relevant criteria for its quantitative estimation. Analysis of common data on various effects upon cerebral vessels shows an essential difference between the functional stability and the idea of "autoregulation of the cerebral circulation". The data obtained suggest that the essential part of the mechanism for active responses of cerebral vessels which maintains the functional stability of this portion of the vascular system, consists of a neurogenic component involving central nervous structures localized, for instance, in the medulla oblongata.

Animals

30. Common problems of current data systems for long-term care.

This article reviews the varied long-term care data systems now being developed in the United States in the framework of a matrix derived from systems analysis and systems theory. Some common problems that emerge are that incentives are coming largely from the societal level; feedback to the institutional and patient care levels is lacking; little attention is being given to the individual level of decision making at one end of the spectrum and the national policy level at the other; and ideas and procedures developed in the acute care setting are being transferred to services traditionally lower in resources and determined more often by levels of patient functioning than by disease. The dominant instrument that emerges in long-term care is the periodic assessment form, in contrast to a hospital discharge abstract or ambulatory care encounter form; however, the data requirements appear to be more voluminous than in the case of acute hospital care, although less manpower is available to respond.

Information Systems

QCatch: a framework for quality control assessment and analysis of single-cell sequencing data.

MOTIVATION: Single-cell sequencing data analysis requires robust quality control (QC) to mitigate technical artifacts and ensure reliable downstream results. While tools like alevin-fry and simpleaf (and augmented execution context for the alevin-fry), offer flexibility and computational efficiency to process single-cell data, this ecosystem will further benefit from a standardized QC reporting tailored for its outputs. RESULTS: We introduce QCatch, a Python-based command-line tool that generates comprehensive and interactive HTML QC reports designed specifically for single-cell quantification results. Taking the output directory of alevin-fry or simpleaf as the input, QCatch is able to perform essential processing steps, like cell calling, and generate detailed QC reports that contain informative visualizations and statistics, including unique molecular identifier (UMI) count distributions, sequencing saturation estimates, and splicing status information, for QC assurance. Built for seamless integration into downstream analysis workflows, QCatch exports the processed results in a richly-annotated H5AD format file, a widely used data format common among many downstream single-cell data analysis tools. AVAILABILITY AND IMPLEMENTATION: The source code and documentation of QCatch are available on GitHub at https://github.com/COMBINE-lab/QCatch. QCatch can be installed via both Bioconda and PyPI.

Single-Cell Analysis

The Data Distillery: A Graph Framework for Semantic Integration and Querying of Biomedical Data.

The Data Distillery Knowledge Graph (DDKG) is a framework for semantic integration and querying of biomedical data across domains. Built for the NIH Common Fund Data Ecosystem, it supports translational research by linking clinical and experimental datasets in a unified graph model. Clinical standards such as ICD-10, SNOMED, and DrugBank are integrated through UMLS, while genomics and basic science data are structured using ontologies and standards such as HPO, GENCODE, Ensembl, STRING, and ClinVar. The DDKG uses a property graph architecture based on the UBKG infrastructure and supports ontology-based ingestion, identifier normalization, and graph-native querying. The system is modular and can be extended with new datasets or schema modules. We demonstrate its utility for informatics queries across eight use cases, including regulatory variant analysis, tissue-specific expression, biomarker discovery, and cross-species variant prioritization. The DDKG is accessible via a public interface, a programmatic API, and downloadable builds for local use.

Journal Article

Excretion of common neutral steroids in healthy subjects as estimated by multi-column chromatography.

Excretion data for common neutral urinary steroids from a total of 330 healthy subjects from different parts of the world and of different sex and age are given. The estimations, which have been performed by multi-column liquid chromatography, include 24 h excretion values for both common 17-oxosteroids and the common metabolites of cortisol, including the cortolones and the cortols. Comparisons are made with values from the world literature and with isotope experiments.

17-Ketosteroids

Permutation tests to assess sex differences in omics data.

It is common to sex-stratify analyses of omics data and to report effects as 'sex-specific' when they are significant in only one sex. However, when analysing hundreds or thousands of molecules, this approach will yield many spurious 'sex-specific' effects if not supported by significant interactions. I illustrate this problem using an RNA sequencing dataset showing almost no significant sex by treatment interactions, but where sex-stratified analyses yield hundreds of 'sex-specific' effects of treatment. These 'sex-specific' effects could be spurious or could be real but not show interactions due to low statistical power. To distinguish these possibilities, I describe permutation tests, which provide an intuitive way to determine if a pattern of observations differs from what would be expected due to chance. For this dataset, assigning sex at random often generates more 'sex-specific' effects than the real data, demonstrating that there is little evidence of sex differences. Next, I simulate an RNA sequencing dataset that includes genes modelled to have sex-specific effects of a condition. As expected, analysis of this simulated dataset yields both significant interactions and sex-specific effects in sex-stratified analyses. While stratified analyses detect a higher number of sex-specific effects than the analysis of interactions, they erroneously identify genes not modelled to show sex-specific effects more often than interactions. A permutation test confirms that the number of sex-specific effects observed in the simulated dataset is greater than expected due to chance. Permutation tests can be applied to omics studies of sex differences, simultaneously providing (i) a clear and simple demonstration of the problems of sex-stratified analyses, and (ii) additional evidence of sex-specific effects where these are present. R code is provided for permutations, simulations, and plots to visualize potential sex-specific effects, which can be adapted to other types of data.

Female

SimpleMicrobiome: An integrated web-based platform for streamlined microbiome data analysis and visualization.

Microbiome studies require multiple analytical steps after initial sequence processing. These steps commonly include data harmonization, preprocessing, taxonomic profiling, diversity analysis, differential abundance testing, predictive modeling, network inference, and preparation of publication-ready outputs. Although robust packages are available for many of these tasks, routine use often depends on command-line workflows, repeated data reformatting, and method-specific scripting. These requirements can limit accessibility for experimental researchers and complicate consistent analysis across interdisciplinary teams. We developed SimpleMicrobiome, a web-based R Shiny platform that integrates established microbiome analysis methods into a single interactive downstream workflow. The application accepts standard abundance, taxonomy, and metadata tables, supports interactive preprocessing and sample filtering, and provides modules for taxa profile visualization, alpha and beta diversity analysis, ANCOM-BC2 and MaAsLin2 differential abundance testing, Random Forest modeling with SHAP-based interpretation, microbial association network inference using SparCC and SPIEC-EASI through NetCoMi, correlation heatmaps, and dbRDA/CAP-style association biplots. The platform is implemented as a modular Shiny application so that preprocessing choices are propagated across downstream analyses, results can be exported as figures and tables, and the same application can be run through the public server, source-code installation, or a Docker image. SimpleMicrobiome consolidates major downstream microbiome analysis tasks in an accessible browser-based environment while retaining links to established analytical frameworks. The platform may reduce technical barriers for non-programming users, improve consistency across exploratory and reporting-oriented analyses, and support collaborative microbiome research. The public application is available at https://simplemicrobiome.mglab.org, the source code is available at https://github.com/yjcho2252/SimpleMicrobiome, and a Docker image for local deployment is available at https://hub.docker.com/r/mglab2252/simplemicrobiome.

differential abundance

Possible effects of non-HLA antibodies in common typing sera on HLA antigen frequency data in leukemia.

Selective adsorption of several common "monospecific" HLA typing sera with HLA typed platelets, purified B lymphocytes, cultured lymphoid cells, or lymphocytes from patients with active chronic lymphocytic leukemia demonstrated that many of these sera contain antibodies to non-HLA antigens. Antibodies were detected to antigens present on peripheral blood B-lymphocytes, cultured lymphoid cells and leukemic cells from patients with both myelocytic and lymphocytic forms of leukemia but absent from T-lymphocytes and platelets. Since these kinds of antibodies appear to be present in a large proportion of common HLA typing sera, caution should be used in interpreting all data related to HLA antigen expression in leukemia.

Antibodies

The role of the interview in student selection.

Four types of data are commonly considered in selecting applicants for academic programs: 1) test scores; 2) grade-point averages; 3) personal data; and 4) interview results. Limitations and advantages of using interview data in the selection decision are discussed. Fifteen suggestions are offered for improving the validity and reliability of interview data. Strategies for data collection and analysis are suggested for validation studies of selection interviews.

Education, Medical

Evidence of a specific nidation site in ruminants.

The site of umbilical cord attachment in ruminants indicates the limited segment of the uterus where the blastocyst attachment occurs and could have potential significance for locating presumptive nidation sites. Measurements of the site of cord attachment were made on impala (Aepyceros melampus) and common duiker (Sylvicapra grimmia) at several stages of gestation. Both implant only in the right uterine horn although they ovulate from either ovary. Relative to uterine length, cord attachment in impala is somewhat closer to the cervix than it is in common duiker. As pregnancy advances in common duiker, the relative position of cord attachment becomes closer to the tubal end. This relationship was not seen in impala and may perhaps to be attributed inadequate data. Upon extrapolation of the data from common duiker, a presumptive attachment area is suggested for this species. This region is located at about 41% of the distance from the internal cervical os to the uterotubal junction. Similar cord attachment data could be used in any ruminant species to indicate the existence and location of a specific nidation site.

Animals

An introductory practical guide to secondary data analysis in pediatric urology.

INTRODUCTION: Secondary data analysis (SDA) has become an increasingly important approach in pediatric urology, enabling the study of long-term outcomes, care variation, and disparities in populations with chronic or congenital urologic conditions. With the growing availability of large datasets, a structured approach to designing and conducting SDA studies is increasingly relevant. OBJECTIVES: To provide an introductory, practical guide to SDA in pediatric urology by (1) summarizing commonly used data sources with representative studies, (2) outlining a stepwise approach to designing and executing SDA studies, and (3) highlighting key methodological considerations, limitations, and opportunities for future work. STUDY DESIGN: Narrative review of existing literature and commonly used datasets relevant to pediatric urology, including administrative claims, hospital encounter databases, clinical registries, electronic health record networks, and population-based surveys. RESULTS: Data sources differ in scope, clinical granularity, longitudinal follow-up, and representativeness, and each is suited to specific research questions. We present a practical workflow for SDA, including dataset selection, cohort definition, and analytic planning. Linkage across datasets can provide a more comprehensive view of care patterns and outcomes, although feasibility is influenced by legal, technical, and data-quality constraints. DISCUSSION: SDA enables population-level analyses and the study of rare conditions that are challenging to evaluate through single-center or prospective designs. However, careful cohort definition, feasibility assessment, and awareness of data limitations are essential to ensure validity and interpretability. CONCLUSION: SDA provides a scalable, cost-efficient framework for generating meaningful evidence in pediatric urology. Continued efforts to harmonize data elements, improve linkage infrastructure, and support cross-institution collaboration will enhance the quality and impact of future research. This article provides a practical framework and examples to support the design and execution of SDA studies.

Humans

Biocomputational methodology an adjunct to theory and applications.

The role of "methodology", as distinguished from "theory" and "application", is discussed and illustrated. It is argued that research in the biomedical sciences is moving towards a degree of complexity different in both kind and extent from that usually encountered in other disciplines. Certain biocomputational methodology can be viewed as a bridge between the data forms commonly encountered in biomedicine, and the statistical and computational machinery which had previously been developed to deal with physical science information. Examples are given of three promising research subareas, all of which concern methods for dealing with highly complex forms of health and medical data.

Computers