PubMed HealthSearch

SEARCH · PubMed Health

Results for “data repository”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository.

Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata. Here, we introduce scBaseCount, a database that leverages an AI agent to automate discovery and metadata extraction and standardize data processing. Built by mining all 10x Genomics datasets, scBaseCount is the largest public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues. It offers an unbiased view of the data landscape within the SRA and enables the training of more performant computational models through access to broader phenotypic diversity. Uniform processing enables measurement of both intronic and exonic reads and non-coding gene expression and improves alignment across experiments. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to autonomously curate biological data repositories.

Single-Cell Analysis

Gencube: centralized retrieval and integration of multi-omics resources from leading databases.

MOTIVATION: The volume of multi-omics data for diverse species is growing at an unprecedented rate, with new genome assemblies, related annotations, and high-throughput sequencing resources being submitted daily to various genomic data repositories. In response to this data influx, both existing and new databases are establishing optimized hierarchical structures to manage the vast amount of information. However, the lack of accessible command-line tools, combined with the functional limitations and unintuitive design of existing options, presents significant challenges for researchers. This gap underscores a critical need for a tool that enables streamlined retrieval and integration of omics data across these diverse repositories. RESULTS: We have developed Gencube, a command-line tool that enables centralized retrieval and integration of a comprehensive set of six different data types-genome assemblies, gene sets, annotations, sequences, comparative genomic data, and NGS-based omics resources-from various leading databases. AVAILABILITY AND IMPLEMENTATION: Gencube is a free and open-source tool, with its code available on GitHub: https://github.com/snu-cdrc/gencube and also archived on Zenodo: https://doi.org/10.5281/zenodo.14607649.

Databases, Genetic

Network-based integration of metabolomics data from large-scale repositories.

INTRODUCTION: Public metabolomics data repositories such as MetaboLights and Metabolomics Workbench host rapidly growing volumes of raw data, processed results, and metadata. As data deposition becomes a prerequisite for funding and publication, there is an increasing need for tools that enable integration and joint reanalysis of datasets across studies to maximise reuse and reproducibility. OBJECTIVES: This study aims to enable large-scale integrative meta-analysis of public metabolomics data, exploiting harmonised metabolite annotations to identify robust multi-study metabolite and pathway signatures and to provide global visual overviews of repository content. METHODS: We developed a network-based integration framework operating at both the study (dataset) level and the metabolite or pathway level. Metabolite-level meta-networks integrate studies with shared biological context using co-occurrences of differential metabolites represented as bipartite graphs. Study-level networks compare observed metabolites for overall repository exploration. Networks can be explored interactively using a dedicated Python Dash app available at https://github.com/EloisaRL/Metabolomic-data-analysis-app/tree/main . RESULTS: As an example, the approach was applied to six COVID-19 plasma datasets from MetaboLights generated using LC-MS and NMR. Ten metabolites were identified as differential in at least three studies, including consistently up-regulated pyroglutamic acid, in agreement with the literature. Pathway-level networks provided an overview of shared biological processes across studies. A global network of 1,181 studies in Metabolomics Workbench demonstrated clustering by assay coverage and associated metadata, as expected. CONCLUSION: Network-based integration of harmonised metabolomics data enables robust cross-study analyses and highlights the critical importance of standardised annotation pipelines. Such approaches enhance the reuse, reproducibility, and impact of public metabolomics datasets, accelerating biological discovery.

Metabolomics

Clostridium difficile Infections after Blunt Trauma: A Different Patient Population?

BACKGROUND: The epidemiology of Clostridium difficile-associated infection (CDI) has changed, and it is evident that susceptibility is related not only to exposures and bacterial potency, but host factors as well. Several small studies have suggested that CDI after trauma is associated with a different patient phenotype. The purpose of this study was to examine and describe the epidemiologic factors associated with C. difficile in blunt trauma patients without traumatic brain injury using the Trauma-Related Database as a part of the "Inflammation and Host Response to Injury" (Glue Grant) and the University of Florida Integrated Data Repository. METHODS: Previously recorded baseline characteristics, clinical data, and outcomes were compared between groups (67 C. difficile and 384 uncomplicated, 813 intermediate, and 761 complicated non-C. difficile patients) as defined by the Glue Grant on admission and at days seven and 14. RESULTS: The majority of CDI patients experienced complicated or intermediate clinical courses. The mean ages of all cohorts were less than 65 y and CDI patients were significantly older than uncomplicated patients without CDI. The CDI patients had increased days in the hospital and on the ventilator, as well as significantly higher new injury severity scores (NISS), and a greater percentage of patients with NISS >34 points compared with non-CDI patients. They also had greater Marshall and Denver multiple organ dysfunction scores than non-CDI uncomplicated patients, and greater creatinine, alkaline phosphatase, neutrophil count, lactic acid, and PiO2:FiO2 compared with all non-CDI cohorts on admission. In addition, the CDI patients had higher glucose concentrations and base deficit from uncomplicated patients and greater leukocytosis than complicated patients on admission. Several of these changes persisted to days seven and 14. CONCLUSION: Analysis of severe blunt trauma patients with C. difficile, as compared with non-CDI patients, reveals evidence of increased inflammation, immunosuppression, worse acute kidney injury, higher NISS, greater days in the hospital and on the ventilator, higher organ injury scores, and prolonged clinical courses. This supports reports of an increased prevalence of CDI in a younger population not believed previously to be at risk. This unique population may have specific genomic or inflammation-related risk factors that may play more important roles in disease susceptibility. Prospective analysis may allow early identification of at-risk patients, creation of novel therapeutics, and improved understanding of how and why C. difficile colonization transforms into infection after severe blunt trauma.

Adolescent

Integrating Biobanking Into Conservation Practice: The Development and Impact of the EAZA Biobank.

Zoological biobanks are becoming essential tools in conservation, offering a means to preserve genetic material and support in situ population management amid accelerating biodiversity loss. With rapid advances in genomics, cryopreservation, and assisted reproduction technologies, biobanks enable a proactive approach to providing insurance against genetic erosion and facilitating future research, supplementation, and genetic rescue. However, to be effective, zoological biobanks must be purposefully designed, strategically integrated into conservation frameworks such as the Convention on Biological Diversity (CBD) Kunming-Montreal Global Biodiversity Framework (KMGBF), and regularly evaluated for coverage and impact. Using the EAZA Biobank as an example, we outline the structure, development, and collaborative foundations that have enabled its rapid growth, built on community support and conservation impact. Leveraging EAZA's institutional network and data-sharing platforms such as ZIMS, the Biobank employs a decentralized, four-hub model of zoological institutions storing samples. A gap analysis, integrating threat status, breeding programs, genomic data repositories, and phylogenetic diversity, highlights current sampling strengths and deficiencies and guides future collection priorities. The integration of specimen-specific genomic data and the EAZA Biobank Cryonetwork of institutions with expertise in storing and generating gametes and cell lines will expand the Biobank's role in population management and conservation. Zoological biobanks must now evolve alongside advances in biotechnology and genomics. Sample collection strategies should serve conservation needs and anticipate future applications in genomics, cryobiology, and conservation medicine, linking biospecimens with the wealth of data generated from them. This approach should be scalable beyond EAZA, forming the foundation of a global standardized biobanking framework. Ultimately, zoological biobanks are not merely repositories of the past-they are essential infrastructures shaping the future potential of species conservation.

EAZA

A consensus guide to preclinical indirect calorimetry experiments.

Understanding the complex factors influencing mammalian metabolism and body weight homeostasis is a long-standing challenge requiring knowledge of energy intake, absorption and expenditure. Using measurements of respiratory gas exchange, indirect calorimetry can provide non-invasive estimates of whole-body energy expenditure. However, inconsistent measurement units and flawed data normalization methods have slowed progress in this field. This guide aims to establish consensus standards to unify indirect calorimetry experiments and their analysis for more consistent, meaningful and reproducible results. By establishing community-driven standards, we hope to facilitate data comparison across research datasets. This advance will allow the creation of an in-depth, machine-readable data repository built on shared standards. This overdue initiative stands to markedly improve the accuracy and depth of efforts to interrogate mammalian metabolism. Data sharing according to established best practices will also accelerate the translation of basic findings into clinical applications for metabolic diseases afflicting global populations.

Calorimetry, Indirect

A decentralized future for the open-science databases.

The continuous and reliable open access to curated biological data repositories is indispensable for accelerating rigorous scientific inquiry and fostering reproducible research outcomes. However, the current paradigm, which relies heavily on centralized infrastructure for the storage and distribution of foundational biomedical datasets, inherently introduces significant vulnerabilities. This centralized model is susceptible to single points of failure, including cyberattacks, technical malfunctions, natural disasters, and even political or funding uncertainties. Such disruptions can lead to widespread data unavailability, data loss, integrity compromises, and substantial delays in critical research, ultimately impeding scientific progress. The downstream effect of such interruptions can be the widespread paralysis of diverse research activities, including computational, clinical, molecular, and climate studies. This scenario vividly illustrates the inherent dangers of consolidating essential scientific resources within a single geopolitical or institutional locus. As data generation is accelerating and the global landscape continues to fluctuate, the sustainability of centralized models must be critically re-evaluated. A shift toward federated and decentralized architectures may offer a robust and forward-looking approach to enhancing the resilience of scientific data infrastructures by reducing exposure to governance instability, infrastructural fragility, and funding volatility, while also promoting equity and global accessibility. Inspired by established models such as ELIXIR's federated infrastructure and the policy and funding frameworks developed by CODATA and the Global Biodata Coalition (GBC), emerging Decentralized Science (DeSci) initiatives can contribute to building more resilient, fair, and incentive-aligned data ecosystems. The future of open science depends on integrating these complementary approaches to establish a globally distributed, economically sustainable, and institutionally robust infrastructure that safeguards scientific data as a public good, further ensuring continued accessibility, interoperability, and preservation for generations to come. Here, we examine the structural limitations of centralized repositories, evaluate federated and decentralized models, and propose a hybrid framework for resilient, fair, and sustainable scientific data stewardship.

data accessibility

WormBase as an integrated platform for the C. elegans ORFeome.

The ORFeome project has validated and corrected a large number of predicted gene models in the nematode C. elegans, and has provided an enormous resource for proteome-scale studies. To make the resource useful to the research and teaching community, it needs to be integrated with other large-scale data sets, including the C. elegans genome, cell lineage, neurological wiring diagram, transcriptome, and gene expression map. This integration is also critical because the ORFeome data sets, like other 'omics' data sets, have significant false-positive and false-negative rates, and comparison to related data is necessary to make confidence judgments in any given data point. WormBase, the central data repository for information about C. elegans and related nematodes, provides such a platform for integration. In this report, we will describe how C. elegans ORFeome data are deposited in the database, how they are used to correct gene models, how they are integrated and displayed in the context of other data sets at the WormBase Web site, and how WormBase establishes connection with the reagent-based resources at the ORFeome project Web site.

Animals

Development of a Blockchain-Based Platform to Enable Indigenous Data Sovereignty and Shared Research Participation With Indigenous Communities: Technology Prototyping and Community Engagement Study.

BACKGROUND: Historic and ongoing problematic practices regarding the collection, storage, and use of Indigenous health data have led to the need to ensure principles of Indigenous Data Sovereignty (IDS) are followed in research practices and technology development. OBJECTIVE: This project, a partnership between UC San Diego and the Native BioData Consortium (NativeBio), sought to explore the practical application of blockchain technology and its potential to facilitate Indigenous-led research collaboration. METHODS: This project first undertook purposeful relationship building with NativeBio to form a Community Advisory Board (CAB) for identifying community and technology needs for a blockchain research collaboration platform with an initial focus on genomic data. Over a 2-year project period, a series of public meetings and presentations at Indigenous-led conferences introduced the concept of exploring compatibility between blockchain and IDS principles, followed by iterative prototyping and co-design of a blockchain platform with NativeBio, using Ethereum as the underlying protocol. RESULTS: Direct engagement with NativeBio and the CAB informed the initial design and development of a "b-IDS" proof-of-concept (POC) blockchain platform. The POC consists of three main components: (1) the web front-end layer, (2) the Ethereum network that executes the smart contract and blockchain storage aspects of the framework, and (3) the back-end database that stores off-chain interactions and data for future use with external genomic data repositories. After refinement of the POC, a community-based participatory research (CBPR) use case aligned with IDS principles was identified as a practical workflow and incorporated into the design of the POC for implementation. CONCLUSIONS: The findings from this project demonstrated the potential use of operationalizing IDS through blockchain technology with proactive and sustained engagement with Indigenous partners. Blockchain technology may have certain advantages over other data governance approaches and systems, facilitating timely oversight, shared decision-making and consent structures, and direct involvement of Indigenous communities in technology design, respecting the core principles of IDS and CBPR. Future development of the blockchain-IDS POC will need to incorporate other research practices and ethics frameworks to expand its use to other public health and biomedical research use cases.

Blockchain

DORSSAA: Drug-Target interactOmics Resource Based on Stability/Solubility Alteration Assay.

Advancements in high-throughput techniques such as Thermal Proteome Profiling and the high-throughput Proteome Integral Solubility Alteration assay have revolutionized our understanding of drug-protein interactions. Despite these innovations, the absence of an integrative platform for cross-study analysis of stability and solubility alteration data represents a significant bottleneck. To address this gap, we introduce Drug-target interactOmics Resource based on Stability/Solubility Alteration Assay (DORSSAA), an interactive and expandable web-based platform for the systematic analysis and visualization of proteome stability and solubility alteration assay datasets. Currently, DORSSAA features 1,135,985 records spanning 38 cell lines and organisms, 135 compounds, and 40,742 protein targets. Through its user-friendly interface, the resource supports comparative drug-protein interaction analysis and facilitates the discovery of actionable therapeutic targets. Through two case studies, methotrexate target profiling in A549 cells and combinatorial-therapy drug-target interactions in leukemia cell lines, we demonstrate DORSSAA's utility for identifying protein-drug interactions across diverse experimental contexts. This resource empowers researchers to accelerate drug discovery and enhance our understanding of protein behavior. Compared with data repositories and interaction databases, DORSSAA provides direct protein-level evidence of mechanisms of action with strict statistical control for each study. This enables more reliable identification of drug targets, off-target effects, and potential drug combinations.

Humans

circASbase: A Comprehensive Database of Alternative Splicing Events in circRNAs.

Although extensive evidence has underscored the critical role of alternative splicing (AS) in generating mature circular RNA (circRNA) isoforms and augmenting their functional diversity, a significant gap remains in the availability of specialized databases housing circRNA AS events. To bridge this gap, we develop circASbase, a pioneering and comprehensive database that catalogs 452,129 AS events in 884,047 full-length circRNAs from 581 samples across 13 species, and provides rich annotations to facilitate understanding the splicing regulation of circRNA. Our findings reveal substantial differences between circRNAs and linear transcripts regarding the distribution and occurrence of AS events, highlighting the unique regulatory landscape of circRNAs. These special splicing events result in functional differences of circRNAs by affecting internal ribosome entry sites, N6-methyladenosine sites, open reading frames, protein features, microRNA targets, and more. In summary, circASbase not only meets the urgent need of the research community for data repositories, but also represents a significant advancement in our understanding of circRNA biology. With its user-friendly interfaces and web-based visualization tools, circASbase is poised to become an indispensable resource for researchers exploring the regulatory mechanisms and functional roles of AS events in circRNAs. This database will continuously drive new insights and discoveries in the field, setting the stage for further advancements in circRNA research. circASbase is freely available at http://reprod.njmu.edu.cn/cgi-bin/circASbase/.

Alternative Splicing

Detectable C-Peptide and Diabetic Ketoacidosis Risk in Type 1 Diabetes.

OBJECTIVE: To investigate whether detectable C-peptide levels in type 1 diabetes is associated with a lower risk of diabetic ketoacidosis (DKA). RESEARCH DESIGN AND METHODS: We analyzed the Diabetes Control and Complications Trial publicly available data repository for the association between detectable stimulated C-peptide (>0.2 nmol/mL, measured annually) and DKA incidence over an average 6.5-year follow-up. We used crude and adjusted Andersen-Gill models for recurrent DKA events with time-dependent covariates. RESULTS: Of the 1,441 participants (53% male, median age 27 years), 129 (9%) experienced 180 DKA events. Of these events, 179 (99.44%) occurred after C-peptide was ≤0.2 nmol/L in the prior year and only 1 event occurred with C-peptide >0.2 nmol/L. C-peptide >0.2 nmol/L was associated with a DKA hazard ratio of 0.07 (95% CI 0.01-0.48; P = 0.007), consistent across adjusted models. CONCLUSIONS: Endogenous insulin production was significantly associated with lower DKA risk, suggesting that treatments preserving insulin production could decrease long-term DKA risk.

C-Peptide

Consistently processed RNA sequencing data from 50 sources enriched for pediatric data.

Larger cohorts improve the power of tumor gene expression analysis, but the signal is muddied if datasets are processed using different methods or have inaccurate metadata. Here we present five compendia containing consistently processed gene expression data derived from 16,446 diverse RNA sequencing datasets. To create the compendia, we obtained access to RNA sequence data from repositories containing public data as well as clinical partners with access to non-published data. We then assessed the quality, quantified gene expression, harmonized clinical metadata, and released the expression values and metadata without access restrictions. These datasets have been used for diverse projects ranging from identifying similarities between tumor types to assessing how well cell lines recapitulate tumors. They have also been used for n-of-1 analysis to identify genes with unusual expression patterns in a single sample and to infer molecular diagnosis. The comparison to new data is enabled by our dockerized, freely available pipeline. The compendia have been cited in at least 20 publications.

Humans

GICPIdb: an archival repository of multimodal data focusing on pathological images for gastrointestinal cancers.

INTRODUCTION: Deep learning (DL) shows great potential for predicting biomarkers from routine histopathological slides of gastrointestinal (GI) cancers. Yet most existing models are validated on limited patient cohorts, while pathological image annotation and molecular marker standardization demand substantial professional expertise. To address these gaps, we constructed the Gastrointestinal Cancer Pathological Image Archive (GICPIdb, gicpidb.shubuzuo.top), a dedicated database and web platform covering seven major GI cancer types. METHODS: High-quality hematoxylin and eosin (H&E)-stained whole-slide images were collected from multiple sources and uniformly processed. Image annotations were performed by board-certified pathologists following standardized protocols. GICPIdb offers five interactive web modules for data uploading, quality control, feature extraction, online annotation and AI-based prediction. Its intuitive interface supports data browsing, retrieval, visualization and downloading. RESULTS: The database houses 2,863 pathologist-annotated, uniformly processed, high-quality H&E stained images collected from 2,655 patients. Of these, 1,699 patients were sourced from The Cancer Genome Atlas (TCGA), 182 from the Clinical Proteomic Tumor Analysis Consortium (CPTAC), and 424 from China-Japan Friendship Hospital and 350 from Chifeng Municipal Hospital in Inner Mongolia, China. It also integrates data on over 50 key molecular markers (e.g., MSI, TMB) and prognostic labels related to survival, recurrence and metastasis. DISCUSSION: GICPIdb aims to promote the development of DL-driven AI tools for cancer research and clinical translation. The multi-institutional data collection and standardized annotation pipeline are expected to enhance the generalizability and reproducibility of AI-based prediction models across diverse patient populations.

deep learning

Computational metabolomics at scale: from open data to insight.

Metabolomics data are currently generated at scale thanks to the evolution of technologies that have led to marked improvements in the number of metabolites detected, spanning all chemical classes. These data are increasingly submitted to public repositories for data reuse, integration, and interpretation. Despite the availability of public resources and associated computational tools, the field still lacks a widely adopted, consistent data and analytics infrastructure capable of transforming this wealth of information into scientific insight. Indeed, the metabolomics field is just now scratching the surface of being able to harness the power of new computational technologies. In this review, we summarize discussions from the "Dagstuhl-Seminar 24181 Computational Metabolomics: Towards Molecules, Models, and their Meaning" with a focus on public data availability, open data standards, data and knowledge integration, and education. Our goal is to raise awareness and adoption of the latest open science resources while highlighting key areas needing further development.

Metabolomics

SeqUIaSCOPE: multi-omics data integration platform for single-patient clinical oncology pathway exploration.

SUMMARY: SeqUIaSCOPE is an open-source platform designed for routine clinical oncology diagnostics through case-centric integration and visualization of genomic variants, fusion events, and expression profiles. The platform combines molecular-level validation via embedded genome browsing with systems-level interpretation through dynamic pathway visualization, enabling geneticists to assess how alterations converge across biological networks. Flexible reporting with customizable templates accommodates diverse institutional requirements, while secure cluster-based or local deployment ensures compliance with data protection policies, making advanced multi-omics diagnostics accessible to academic and clinical institutions. AVAILABILITY AND IMPLEMENTATION: SeqUIaSCOPE is freely available on GitHub at https://github.com/BioIT-CEITEC/sequiascope under the MIT license and archived at Zenodo (https://zenodo.org/records/21338445). Due to the sensitive nature of patient data, the repository provides simulated datasets that mimic the structure of real clinical data for testing and exploration. Documentation and a live demo accompany these datasets, allowing users to explore the application without any prior setup. The repository also includes a Helm chart for Kubernetes deployment and Docker containers for local deployment, ensuring compatibility across Linux, macOS, and Windows. No user registration is required, and all data remains on local or institutional infrastructure.

Humans

Integrative analysis of single-cell sequencing identifies CD8+ TIM3+ CD101+ T cell-associated genes as prognostic biomarkers in breast cancer.

BACKGROUND: Breast cancer is a prevalent and deadly malignancy that significantly impacts women's quality of life and imposes financial burdens. Despite therapeutic advancements, tumour heterogeneity and frequent relapses remain major challenges. Accordingly, this study aimed to characterize immune features associated with CD8+ TIM3+ CD101+ T cells and develop a prognostic signature for breast cancer. METHODS: This study integrated single-cell and bulk transcriptomic datasets to characterize CD8+ TIM3+ CD101+ T cell (CCT)-related immune features and construct a prognostic signature in breast cancer. Single-cell RNA-seq data were sourced from the Gene Expression Omnibus (GEO) repository, and bulk transcriptomic data were from The Cancer Genome Atlas (TCGA) and GEO databases. Analytical methods included pseudo-time trajectory reconstruction (Monocle2), intercellular signalling analysis (CellChat), functional enrichment (ClusterProfiler), and immune profiling (ssGSEA). Prognostic modeling was conducted using least absolute shrinkage and selection operator (LASSO) Cox regression, with validation via Kaplan-Meier and time-dependent receiver operating characteristic (ROC) analyses. RESULTS: Single-cell analysis identified 17 clusters spanning seven cell types, including T cells, myeloid cells, and epithelial cells. T-cell sub-clustering revealed four subtypes. Pseudotime analysis suggested a potential state-transition relationship between CD8+ CD101- TIM3+ and CD8+ CD101+ TIM3+ T-cell states. A total of 121 differentially expressed genes were enriched in vital biological processes. An 11-gene prognostic model showed strong predictive power across cohorts. Single-cell T-cell reclustering identified a CD8+ CD101+ TIM3+ T-cell subpopulation, which was primarily characterized by the expression of markers such as CD101 and HAVCR2/TIM3. CONCLUSIONS: This study maps cellular heterogeneity and molecular networks in breast cancer, offering insights for targeted therapy and improved prognosis.

Breast invasive carcinoma

Freely available genomic datasets for atrial fibrillation research: current resources and analytical pipeline.

Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia, characterized by clinical and genetic heterogeneity. Increasing use of genomics and other omics approaches has driven reliance on publicly available AF datasets to advance biological discovery. Thus, this systematic review aimed to identify freely available genomic AF datasets through Mendeley Data and its interconnected repositories, and to characterize the most common analyses performed on these data. The search was conducted in adherence to the PRISMA 2020 guideline. Nineteen freely available genomic AF datasets were identified: Summary statistics for 'Biobank-driven genomic discovery yields new insight into atrial fibrillation biology', hum0014.v8.58qt.v1, AF GWAS in UK Biobank, UK Biobank (Publication 9659), GWAS summary statistics from a 2025 multi-ancestry AF meta-analysis, GSE115574, GSE128188, GSE14975, GSE2240, GSE238242, GSE254133, GSE261170, GSE271748, GSE271839, GSE293813, GSE294456, GSE31821, GSE41177, and GSE79768. The GEO datasets were further examined using differential gene expression, functional enrichment, protein-protein interaction networks, hub gene analysis, microRNA target prediction, and gene clustering, as well as, for the more recently deposited datasets, eQTL colocalization, single-cell/single-nucleus clustering, cell-cell communication analysis, and gene-dosage-dependent transcriptional and electrophysiological profiling. These analyses show some consistency but also considerable heterogeneity in initial conditions, data normalization, and analytical methodological settings. In conclusion, only a limited number of datasets are freely available, so additional, well-characterized and standardized datasets are needed to provide a complete picture of the AF pathology.

Mendeley Data