PubMed HealthSearch

SEARCH · PubMed Health

Results for “Data Analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A Practical Workflow for Spatial Transcriptomics Data Analysis: From Data Acquisition to Advanced Analyses.

Spatial transcriptomics (ST) profiles genome-wide gene expression while preserving the two-dimensional spatial context of mRNA molecules within tissue sections, enabling studies of tissue architecture and microenvironment-associated biology. However, ST analysis remains challenging because data import, quality control, integration, deconvolution, spatial statistics, and visualization often require multiple software environments and reproducible parameter choices. This protocol presents a practical computational workflow for public ST datasets in R, beginning with data acquisition and software setup and proceeding through Seurat-based data loading, quality control, normalization, multi-sample integration, clustering, and spatially variable gene analysis. The workflow then applies complementary deconvolution strategies, including reference-guided SPOTlight analysis and unsupervised STdeconvolve topic modeling, followed by Giotto-based spatial cell-cell communication analysis and interactive region-of-interest (ROI) selection using a custom Python Dash application. By emphasizing script-based execution, explicit parameter rationales, expected outputs, and troubleshooting checkpoints, the protocol provides an adaptable framework for standard array-based ST datasets and related platforms after dataset- and platform-specific parameter evaluation.

Spatial Transcriptomics

[Graphical methods in data analysis (author's transl)].

Data analysis is concerned with attentive description and communication of the information contents of a body of data. Background information, conceptual insight and especially graphical methods play a key role in data analysis for developing a feeling for the data both by formal procedures to be applied in the light of specified models and even more by informal inference or methods that are suggestive and conctructive. This paper reviews graphical methods useful for description, screening, analysis, cross-examining, selection, reduction, presentation and summary of data: for uncovering distributional peculiarities and understanding the structure underlying experimental and survey data. Moreover scatter plots, probability plots and residual plots provide insight into the possible inappropriateness of certain assumptions of the statistical model. Some techniques are illustrated by examples: four-dimensional data may be reprented as scatter plot on ordinary graph paper by using a combination of 2 different sets of symbols for at most 7 different levels of the third variable (formula: see text) and of the fourth variable (formula: see text). Comments on the use of tables and graphical methods, a small overview of the latter and of the scope of applications endeavour to pave the way such that structures may be better understandable and unanticipated characteristics may be spotted.

Factor Analysis, Statistical

An introductory practical guide to secondary data analysis in pediatric urology.

INTRODUCTION: Secondary data analysis (SDA) has become an increasingly important approach in pediatric urology, enabling the study of long-term outcomes, care variation, and disparities in populations with chronic or congenital urologic conditions. With the growing availability of large datasets, a structured approach to designing and conducting SDA studies is increasingly relevant. OBJECTIVES: To provide an introductory, practical guide to SDA in pediatric urology by (1) summarizing commonly used data sources with representative studies, (2) outlining a stepwise approach to designing and executing SDA studies, and (3) highlighting key methodological considerations, limitations, and opportunities for future work. STUDY DESIGN: Narrative review of existing literature and commonly used datasets relevant to pediatric urology, including administrative claims, hospital encounter databases, clinical registries, electronic health record networks, and population-based surveys. RESULTS: Data sources differ in scope, clinical granularity, longitudinal follow-up, and representativeness, and each is suited to specific research questions. We present a practical workflow for SDA, including dataset selection, cohort definition, and analytic planning. Linkage across datasets can provide a more comprehensive view of care patterns and outcomes, although feasibility is influenced by legal, technical, and data-quality constraints. DISCUSSION: SDA enables population-level analyses and the study of rare conditions that are challenging to evaluate through single-center or prospective designs. However, careful cohort definition, feasibility assessment, and awareness of data limitations are essential to ensure validity and interpretability. CONCLUSION: SDA provides a scalable, cost-efficient framework for generating meaningful evidence in pediatric urology. Continued efforts to harmonize data elements, improve linkage infrastructure, and support cross-institution collaboration will enhance the quality and impact of future research. This article provides a practical framework and examples to support the design and execution of SDA studies.

Humans

A data analysis microcomputer package (DAMP) for biomedical signals.

The advent of cheap, powerful microcomputer systems makes the analysis of data via sophisticated techniques available to the personnel who are non-specialists in computing systems. The DAMP package described here is intended for use on personal computers and has therefore been written in BASIC for portability. The analysis techniques are powerful, comprising algorithms to perform sample-data generation, plotting displays, digital data filtering, auto-correlation functions, fast Fourier transforms and autoregressive modelling. The last technique contains a number of options including the display of z-plane plots, frequency response of the model, residual plotting and auto-correlation of the residuals. Illustrative results are shown from psychological mood data and rat locomotor activity. The package is designed both to instruct a user in the techniques of spectral analysis, and also to provide a range of methods for investigating time and frequency behaviour of biomedical data.

Biomedical Engineering

SpectroPipeR-a streamlining post Spectronaut® DIA-MS data analysis R package.

SUMMARY: Proteome studies frequently encounter challenges in down-stream data analysis due to limited bioinformatics resources, rapid data generation, and variations in analytical methods. To address these issues, we developed SpectroPipeR, an R package designed to streamline data analysis tasks and provide a comprehensive, standardized pipeline for Spectronaut® DIA-MS data. This novel package automates various analytical processes, including XIC plots, ID rate summary, normalization, batch and covariate adjustment, relative protein quantification, multivariate analysis, and statistical analysis, while generating interactive HTML reports for e.g. ELN systems. AVAILABILITY AND IMPLEMENTATION: The SpectroPipeR package (manual: https://stemicha.github.io/SpectroPipeR/) was written in R and is freely available on GitHub (https://github.com/stemicha/SpectroPipeR).

Software

UALCAN Mobile, an app for cancer proteogenomic data analysis.

Cancer is a complex disease affecting various organs and is a major cause of death worldwide. During cancer initiation, disease progression, and tumor metastasis, various genomic and proteomic alterations are observed. Recent technological advances have led to the generation of large amounts of molecular data, including genomics and transcriptomics. These large-scale datasets can be utilized to analyze and identify sub-class-specific cancer biomarkers and targets. However, there is a need for the development of user-friendly tools for large-scale data analysis, disseminating the analyzed data in a visualizable format to cancer researchers with no programming skills. We developed UALCAN, a comprehensive platform that allows users to integrate disparate data to better understand the genes, proteins, and pathways perturbed in cancer and make discoveries of potential biomarkers and targets. In the current study, we describe the development of the UALCAN Mobile application (app) that will provide cancer transcriptomic data obtained from The Cancer Genome Atlas (TCGA) project to evaluate protein-coding gene expression based on various stratifications, including stage, grade, race, gender, and molecular-subtypes across over 30 types of cancers. In addition, the UALCAN mobile provides data analysis options for epigenetic changes due to DNA promoter methylation and Clinical Proteomic Tumor Analysis Consortium (CPTAC) cancer proteomic data. The app provides access to large cancer molecular datasets on the go. To find changes in the expression of causative genes and proteins and to identify biomarkers and therapeutic targets, UALCAN mobile app will be extremely valuable. The "UALCAN Mobile" app is free to use and can be downloaded from both the iOS/Apple and the Android Play Store and has been downloaded over 100 times in each of iOS and android app stores.

app

BioNeuralNet: a graph neural network based Multi-Omics network data analysis tool.

SUMMARY: Multi-omics data offer unprecedented insights into complex biological systems, yet their high dimensionality, sparsity, and intricate interactions pose significant analytical challenges. Network-based approaches have advanced multi-omics research by effectively capturing biologically relevant relationships among molecular features (e.g., genes, proteins, metabolites). While these methods are powerful for representing molecular interactions, there remains a need for tools specifically designed to effectively utilize these network representations across diverse downstream analyses. To fulfill this need, we introduce BioNeuralNet, a flexible and modular Python framework tailored for end-to-end network-based multi-omics data analysis. BioNeuralNet leverages Graph Neural Networks (GNNs) to learn biologically meaningful low-dimensional representations from multi-omics networks, converting these complex molecular networks into versatile embeddings. BioNeuralNet supports all major stages of multi-omics network analysis, including several network construction techniques, generation of low-dimensional representations, and a broad range of downstream analytical tasks. Its extensive utilities, including diverse GNN architectures, and compatibility with established Python packages (e.g., scikit-learn, PyTorch, NetworkX), enhance usability and facilitate quick adoption. BioNeuralNet is an open-source, user-friendly, and extensively documented framework designed to support flexible and reproducible multi-omics network analysis in precision medicine. AVAILABILITY AND IMPLEMENTATION: The BioNeuralNet library is available via The Python Package Index (PyPI). Source code, documentation, tutorials, and workflows are hosted at https://bioneuralnet.readthedocs.io. Code archived at https://doi.org/10.5281/zenodo.17503083.

Graph Neural Networks

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score = 0.67-0.90) in genus diversity and showed a high correlation (rSpearman = 0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics

DeeDeeExperiment: building an infrastructure for integrating and managing omics data analysis results in R/Bioconductor.

SUMMARY: Modern omics experiments now involve multiple conditions and complex designs, producing an increasingly large set of differential expression and functional enrichment analysis results. However, no standardized data structure exists to store and contextualize these results together with their metadata, leaving researchers with an unmanageable and potentially non-reproducible collection of results that are difficult to navigate and/or share. Here we introduce DeeDeeExperiment, a new S4 class for managing and storing omics data analysis results, implemented within the Bioconductor ecosystem, which promotes interoperability, reproducibility and good documentation. This class extends the widely used SingleCellExperiment object by introducing dedicated slots for Differential Expression (DEA) and Functional Enrichment Analysis (FEA) results, allowing users to organize, store, and retrieve information on multiple contrasts and associated metadata within a single data object, ultimately streamlining the management and interpretation of many omics datasets. AVAILABILITY AND IMPLEMENTATION: DeeDeeExperiment is available on Bioconductor under the MIT license (https://bioconductor.org/packages/DeeDeeExperiment), with its development version also available on Github (https://github.com/imbeimainz/DeeDeeExperiment).

Software

AWGE-ESPCA: An edge sparse PCA model based on adaptive noise elimination regularization and weighted gene network for Hermetia illucens genomic data analysis.

Hermetia illucens is an important insect resource. Studies have shown that exploring the effects of Cu2+-stressed on the growth and development of the Hermetia illucens genome holds significant scientific importance. There are three major challenges in the current studies of Hermetia illucens genomic data analysis: firstly, the lack of available genomic data which limits researchers in Hermetia illucens genomic data analysis. Secondly, to the best of our knowledge, there are no Artificial Intelligence (AI) feature selection models designed specifically for Hermetia illucens genome. Unlike human genomic data, noise in Hermetia illucens data is a more serious problem. Third, how to choose those genes located in the pathway enrichment region. Existing models assume that each gene probe has the same priori weight. However, researchers usually pay more attention to gene probes which are in the pathway enrichment region. Based on the above challenges, we initially construct experiments and establish a new Cu2+-stressed Hermetia illucens growth genome dataset. Subsequently, we propose AWGE-ESPCA: an edge Sparse PCA model based on adaptive noise elimination regularization and weighted gene network. The AWGE-ESPCA model innovatively proposes an adaptive noise elimination regularization method, effectively addressing the noise challenge in Hermetia illucens genomic data. We also integrate the known gene-pathway quantitative information into the Sparse PCA(SPCA) framework as a priori knowledge, which allows the model to filter out the gene probes in pathway-rich regions as much as possible. Ultimately, this study conducts five independent experiments and compared four latest Sparse PCA models as well as representative supervised and unsupervised baseline models to validate the model performance. The experimental results demonstrate the superior pathway and gene selection capabilities of the AWGE-ESPCA model. Ablation experiments validate the role of the adaptive regularizer and network weighting module. To summarize, this paper presents an innovative unsupervised model for Hermetia illucens genome analysis, which can effectively help researchers identify potential biomarkers. In addition, we also provide a working AWGE - ESPCA model code in the address: https://github.com/yhyresearcher/AWGE_ESPCA.

Animals

Augmented kurtosis-based projection pursuit: a novel, advanced machine learning approach for multi-omics data analysis and integration.

Due to the heterogeneity of multi-omics data, exacting their maximum information potential remains a challenge. Whereas some solutions have been offered, most cannot overcome the large linear dynamic range associated with such data, while others require large biological effect sizes to produce meaningful models. Here, we (i) perform a comprehensive benchmarking of multi-omics data analysis tools, and (ii) introduce kurtosis-based projection pursuit analysis, augmented with classification and regression trees (kPPA-CART) as a robust, easy-to-implement alternative. Using ground truth data, we demonstrate that kPPA-CART exhibits superiority in inferring biological significance from low-intensity (low-count) features and studies with small biological effect sizes. Applying it to experimental breast cancer data from The Cancer Genome Atlas, we identify novel genes that cluster the samples into subtypes that mimic the canonical PAM50 classes with notable improvements. Validating with external metastatic breast cancer data from the AURORA US consortium, kPPA-CART identifies genes that are associated with poor event-free survival and additional clustering associated with increased tumor mutational burden. Finally, we provide an R package and an online implementation of kPPA-CART.

Humans

A hierarchical, count-based model highlights challenges in scATAC-seq data analysis and points to opportunities to extract finer-resolution information.

BACKGROUND: Data from Single-cell Assay for Transposase Accessible Chromatin with Sequencing (scATAC-seq) is highly sparse. While current computational methods feature a range of transformation procedures to extract meaningful information, major challenges remain. RESULTS: Here, we discuss the major scATAC-seq data analysis challenges such as sequencing depth normalization and region-specific biases. We present a hierarchical count model that is motivated by the data generating process of scATAC-seq data. Our simulations show that current scATAC-seq data, while clearly containing physical single-cell resolution, are too sparse to infer true informational-level single-cell, single-region of chromatin accessibility states. CONCLUSIONS: While the broad utility of scATAC-seq at a cell type level is undeniable, describing it as fully resolving chromatin accessibility at single-cell resolution, particularly at individual locus level, may overstate the level of detail currently achievable. We conclude that chromatin accessibility profiling at true single-cell, single-region resolution is challenging with current data sensitivity, but that it may be achieved with promising developments in optimizing the efficiency of scATAC-seq assays.

Single-Cell Analysis

Implementing a training resource for large-scale genomic data analysis in the All of Us Researcher Workbench.

A lack of representation in genomic research and limited access to computational training create barriers for many researchers seeking to analyze large-scale genetic datasets. The All of Us Research Program provides an unprecedented opportunity to address these gaps by offering genomic data from a broad range of participants, but its impact depends on equipping researchers with the necessary skills to use it effectively. The All of Us Biomedical Researcher (BR) Scholars Program at Baylor College of Medicine aims to break down these barriers by providing early-career researchers with hands-on training in computational genomics through the All of Us Evenings with Genetics Research Program. The year-long program begins with the faculty summit, an in-person computational boot camp that introduces scholars to foundational skills for using the All of Us dataset via a cloud-based research environment. The genomics tutorials focus on genome-wide association studies (GWASs), utilizing Jupyter Notebooks and the Hail computing framework to provide an accessible and scalable approach to large-scale data analysis. Scholars engage in hands-on exercises covering data preparation, quality control, association testing, and result interpretation. By the end of the summit, participants will have successfully conducted a GWAS, visualized key findings, and gained confidence in computational resource management. This initiative expands access to genomic research by equipping early-career researchers from a variety of backgrounds with the tools and knowledge to analyze All of Us data. By lowering barriers to entry and promoting the study of representative populations, the program fosters innovation in precision medicine and advances equity in genomic research.

Humans

Analog processing of vestibular nystagmus for on-line cross- correlation data analysis.

An analog processing circuit is described which allow accurate measurement of the phase relationships between input angular acceleration and resulting eye velocity. Vestibular nystagmic data are processed via analog technics to yield slowphase eye velocity. The turntable velocity input is cross-correlated with the eye velocity output, using a Nicolet MED-80 minicomputer system. The resulting correlograms are further processed to obtain precise phase information. Test data analysis shows a system resolution within 1 degree. Data from human and animal subjects are portrayed.

Acceleration

HLA-D typing with lymphoblastoid cell lines. VII. A computer program for data analysis.

When lymphoblastoid cell lines (LCL) are substituted for peripheral blood lymphocytes from human typing cell donors in HLA-D typing experiments, a data analysis program must be designed to distinguish the effect of allo-reactivity from those peculiar to LCL, mainly the "autologous-stimulation" effect. The computer program described in this report was created specifically for such an analysis. The rationale for the design of this program is presented in the preceding report (see this issue).

Cell Line

[Planning and data analysis in prospective controlled clinical trials (author's transl)].

Planning of prospective controlled clinical trials in surgery requires the use of test and control groups, sufficiently frequent repetition of experiments, random allocation of patients to the groups (example), and balancing. The descriptive data analysis should be performed in a stepwise manner (list of new data, rank list, range, median, quartiles, histogram, mean value standard deviation). The advantages of the median-quartile-system and the prerequisites for application of various significance tests are pointed out. In the conduct of controlled clinical trials, the consultative role of experimental surgeons is proposed.

Clinical Trials as Topic

Novelty seeking and rapid symptom improvement across active and sham accelerated iTBS conditions: A pooled individual-patient data analysis.

INTRODUCTION: Major depressive disorder (MDD) is highly prevalent and often treatment-resistant. Accelerated intermittent theta burst stimulation (aiTBS) is a promising intervention for treatment-resistant depression (TRD), though outcomes vary. Personality traits have been examined in relation to rTMS outcomes, yet their role in aiTBS remains underexplored. This pooled individual-patient-data analysis of two randomized, sham-controlled trials examined associations between baseline Temperament and Character Inventory (TCI) traits and one-week symptom change, and whether they differed by condition. METHODS: The left dorsolateral prefrontal cortex was targeted for 20 sessions over 4 days. Personality was assessed with the TCI, depression severity with the 17-item Hamilton Depression Rating Scale (HDRS-17). TCI-symptom-change associations were examined with a robust linear mixed-effects model, adjusting for age, gender, repeated measurements, and study membership. RESULTS: 104 participants were included (M/F 45/59; mean age 40.9 ± 12.7; active/sham 50/54). The model yielded a Time × Novelty Seeking interaction (β = -1.70, p = 0.021): higher baseline Novelty Seeking was associated with faster symptom reduction, without a between-arm difference. However, the interaction did not survive Holm correction across 14 trait-interaction tests (adjusted p = 0.294) and is therefore exploratory. No other interaction reached the uncorrected threshold. CONCLUSIONS: Higher baseline Novelty Seeking showed a nominal association with faster symptom reduction, without a difference between active and sham conditions. Because it did not survive multiplicity correction and was not reproduced in within-arm analyses, it is preliminary and may reflect contextual or nonspecific processes. Independent replication is required before temperament assessment can be clinically informative.

Humans

Physiological consequences of experimental cerebral missile injury and use of data analysis to predict survival.

The authors describe cerebrovascular and cerebral metabolic changes in monkeys, subjected to cerebral missile injury. After injury with BB pellet at 90 m/sec, there is a rapid rise in intracranial pressure (ICP), which reaches a peak 2 to 5 minutes posttrauma, and then falls to about 20 to 30 mm Hg. This, with a fall in mean blood pressure (MBP), results in a 50% reduction in cerebral perfusion pressure (CPP), Cerebral blood flow (CBF) is also reduced, although acutely there is no close relationship with (CPP). Cerebrovascular resistance falls initially and then at 30 minutes rises to very high values. Cerebral metabolic rates (CMR's) for oxygen fall after injury and remain low for the rest of the animal's life; CMR's for lactate rise immediately after injury and persists for 5 hours, then fall. After injury with a faster missile (180 m/sec), the ICP rises higher and faster, and the peak is shorter. The CCP is reduced in this injury to approximately 30 mm Hg, and only one animal survived more than 1 hour. With the conventional forms of data analysis, the length of survival after injury correlates well with MBP, ICP, and CBF, but separately they were completely unsatisfactory for prediction of an individuals prognosis. With the technique of multiple linear regression analysis, the survival of individual animals could be predicted with great accuracy. This is possible also when two postinjury parameters,CBF and MBP, are used.

Animals