PubMed HealthSearch

SEARCH · PubMed Health

Results for “Data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A Practical Workflow for Spatial Transcriptomics Data Analysis: From Data Acquisition to Advanced Analyses.

Spatial transcriptomics (ST) profiles genome-wide gene expression while preserving the two-dimensional spatial context of mRNA molecules within tissue sections, enabling studies of tissue architecture and microenvironment-associated biology. However, ST analysis remains challenging because data import, quality control, integration, deconvolution, spatial statistics, and visualization often require multiple software environments and reproducible parameter choices. This protocol presents a practical computational workflow for public ST datasets in R, beginning with data acquisition and software setup and proceeding through Seurat-based data loading, quality control, normalization, multi-sample integration, clustering, and spatially variable gene analysis. The workflow then applies complementary deconvolution strategies, including reference-guided SPOTlight analysis and unsupervised STdeconvolve topic modeling, followed by Giotto-based spatial cell-cell communication analysis and interactive region-of-interest (ROI) selection using a custom Python Dash application. By emphasizing script-based execution, explicit parameter rationales, expected outputs, and troubleshooting checkpoints, the protocol provides an adaptable framework for standard array-based ST datasets and related platforms after dataset- and platform-specific parameter evaluation.

Spatial Transcriptomics

The molecular function of hemoglobin as reflected in ligand binding data: analysis of data on erythrocytes.

Hemoglobin oxygen binding data on erythrocytes at diffrent pH, PCO2 and bisphosphoglycerate concentrations have been analyzed in terms of an extended version of the Herzfeld-Stanley model of 1972. The binding of oxygen to subunits when the tetramer is in the quaternary oxy conformation was found to be insensitive to moderate changes in pH and pCO2 (0.71 +/- 0.05 mm Hg-1). Utilizing this circumstance it has been possible to obtain, for the first time, unique estimates of energy parameters related to hemoglobin cooperativity and effector action. At 37 degrees C, pH 7.2 and pCO2 22 mm Hg the following parameter values were obtained: The allosteric constant: (1.5 + 0.4)-10(4); the oxygen binding constant of the deoxy state: (5.4 +/- 0.3).10(-3) mm Hg-1; the 2,3-bisphosphoglycerate binding constants: (3.3 +/- 1.3).10(3)1. mol-1(deoxy), (1.3 +/- 0.5).10(2)1. mol-1 (oxy). Quarternary transition most likely takes place after binding of the second O2 molecule. Following the concepts of Perutz the results suggest that (1) protons and carbon dioxide act as constraint effectors and/or as quaternary effectors; (2) the difference in total conformational energy between the two quaternary ligand-free states is almost exclusively confined to molecular constraints and very little to the difference in quaternary conformational energy. The consistency of the results indicate that the model may be regarded as a useful tool for the description of the functional interrelations in the hemoglobin oxygenation process as reflected in oxygen binding data.

Binding Sites

[Graphical methods in data analysis (author's transl)].

Data analysis is concerned with attentive description and communication of the information contents of a body of data. Background information, conceptual insight and especially graphical methods play a key role in data analysis for developing a feeling for the data both by formal procedures to be applied in the light of specified models and even more by informal inference or methods that are suggestive and conctructive. This paper reviews graphical methods useful for description, screening, analysis, cross-examining, selection, reduction, presentation and summary of data: for uncovering distributional peculiarities and understanding the structure underlying experimental and survey data. Moreover scatter plots, probability plots and residual plots provide insight into the possible inappropriateness of certain assumptions of the statistical model. Some techniques are illustrated by examples: four-dimensional data may be reprented as scatter plot on ordinary graph paper by using a combination of 2 different sets of symbols for at most 7 different levels of the third variable (formula: see text) and of the fourth variable (formula: see text). Comments on the use of tables and graphical methods, a small overview of the latter and of the scope of applications endeavour to pave the way such that structures may be better understandable and unanticipated characteristics may be spotted.

Factor Analysis, Statistical

An introductory practical guide to secondary data analysis in pediatric urology.

INTRODUCTION: Secondary data analysis (SDA) has become an increasingly important approach in pediatric urology, enabling the study of long-term outcomes, care variation, and disparities in populations with chronic or congenital urologic conditions. With the growing availability of large datasets, a structured approach to designing and conducting SDA studies is increasingly relevant. OBJECTIVES: To provide an introductory, practical guide to SDA in pediatric urology by (1) summarizing commonly used data sources with representative studies, (2) outlining a stepwise approach to designing and executing SDA studies, and (3) highlighting key methodological considerations, limitations, and opportunities for future work. STUDY DESIGN: Narrative review of existing literature and commonly used datasets relevant to pediatric urology, including administrative claims, hospital encounter databases, clinical registries, electronic health record networks, and population-based surveys. RESULTS: Data sources differ in scope, clinical granularity, longitudinal follow-up, and representativeness, and each is suited to specific research questions. We present a practical workflow for SDA, including dataset selection, cohort definition, and analytic planning. Linkage across datasets can provide a more comprehensive view of care patterns and outcomes, although feasibility is influenced by legal, technical, and data-quality constraints. DISCUSSION: SDA enables population-level analyses and the study of rare conditions that are challenging to evaluate through single-center or prospective designs. However, careful cohort definition, feasibility assessment, and awareness of data limitations are essential to ensure validity and interpretability. CONCLUSION: SDA provides a scalable, cost-efficient framework for generating meaningful evidence in pediatric urology. Continued efforts to harmonize data elements, improve linkage infrastructure, and support cross-institution collaboration will enhance the quality and impact of future research. This article provides a practical framework and examples to support the design and execution of SDA studies.

Humans

A data analysis microcomputer package (DAMP) for biomedical signals.

The advent of cheap, powerful microcomputer systems makes the analysis of data via sophisticated techniques available to the personnel who are non-specialists in computing systems. The DAMP package described here is intended for use on personal computers and has therefore been written in BASIC for portability. The analysis techniques are powerful, comprising algorithms to perform sample-data generation, plotting displays, digital data filtering, auto-correlation functions, fast Fourier transforms and autoregressive modelling. The last technique contains a number of options including the display of z-plane plots, frequency response of the model, residual plotting and auto-correlation of the residuals. Illustrative results are shown from psychological mood data and rat locomotor activity. The package is designed both to instruct a user in the techniques of spectral analysis, and also to provide a range of methods for investigating time and frequency behaviour of biomedical data.

Biomedical Engineering

SpectroPipeR-a streamlining post Spectronaut® DIA-MS data analysis R package.

SUMMARY: Proteome studies frequently encounter challenges in down-stream data analysis due to limited bioinformatics resources, rapid data generation, and variations in analytical methods. To address these issues, we developed SpectroPipeR, an R package designed to streamline data analysis tasks and provide a comprehensive, standardized pipeline for Spectronaut® DIA-MS data. This novel package automates various analytical processes, including XIC plots, ID rate summary, normalization, batch and covariate adjustment, relative protein quantification, multivariate analysis, and statistical analysis, while generating interactive HTML reports for e.g. ELN systems. AVAILABILITY AND IMPLEMENTATION: The SpectroPipeR package (manual: https://stemicha.github.io/SpectroPipeR/) was written in R and is freely available on GitHub (https://github.com/stemicha/SpectroPipeR).

Software

UALCAN Mobile, an app for cancer proteogenomic data analysis.

Cancer is a complex disease affecting various organs and is a major cause of death worldwide. During cancer initiation, disease progression, and tumor metastasis, various genomic and proteomic alterations are observed. Recent technological advances have led to the generation of large amounts of molecular data, including genomics and transcriptomics. These large-scale datasets can be utilized to analyze and identify sub-class-specific cancer biomarkers and targets. However, there is a need for the development of user-friendly tools for large-scale data analysis, disseminating the analyzed data in a visualizable format to cancer researchers with no programming skills. We developed UALCAN, a comprehensive platform that allows users to integrate disparate data to better understand the genes, proteins, and pathways perturbed in cancer and make discoveries of potential biomarkers and targets. In the current study, we describe the development of the UALCAN Mobile application (app) that will provide cancer transcriptomic data obtained from The Cancer Genome Atlas (TCGA) project to evaluate protein-coding gene expression based on various stratifications, including stage, grade, race, gender, and molecular-subtypes across over 30 types of cancers. In addition, the UALCAN mobile provides data analysis options for epigenetic changes due to DNA promoter methylation and Clinical Proteomic Tumor Analysis Consortium (CPTAC) cancer proteomic data. The app provides access to large cancer molecular datasets on the go. To find changes in the expression of causative genes and proteins and to identify biomarkers and therapeutic targets, UALCAN mobile app will be extremely valuable. The "UALCAN Mobile" app is free to use and can be downloaded from both the iOS/Apple and the Android Play Store and has been downloaded over 100 times in each of iOS and android app stores.

app

BioNeuralNet: a graph neural network based Multi-Omics network data analysis tool.

SUMMARY: Multi-omics data offer unprecedented insights into complex biological systems, yet their high dimensionality, sparsity, and intricate interactions pose significant analytical challenges. Network-based approaches have advanced multi-omics research by effectively capturing biologically relevant relationships among molecular features (e.g., genes, proteins, metabolites). While these methods are powerful for representing molecular interactions, there remains a need for tools specifically designed to effectively utilize these network representations across diverse downstream analyses. To fulfill this need, we introduce BioNeuralNet, a flexible and modular Python framework tailored for end-to-end network-based multi-omics data analysis. BioNeuralNet leverages Graph Neural Networks (GNNs) to learn biologically meaningful low-dimensional representations from multi-omics networks, converting these complex molecular networks into versatile embeddings. BioNeuralNet supports all major stages of multi-omics network analysis, including several network construction techniques, generation of low-dimensional representations, and a broad range of downstream analytical tasks. Its extensive utilities, including diverse GNN architectures, and compatibility with established Python packages (e.g., scikit-learn, PyTorch, NetworkX), enhance usability and facilitate quick adoption. BioNeuralNet is an open-source, user-friendly, and extensively documented framework designed to support flexible and reproducible multi-omics network analysis in precision medicine. AVAILABILITY AND IMPLEMENTATION: The BioNeuralNet library is available via The Python Package Index (PyPI). Source code, documentation, tutorials, and workflows are hosted at https://bioneuralnet.readthedocs.io. Code archived at https://doi.org/10.5281/zenodo.17503083.

Graph Neural Networks

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score = 0.67-0.90) in genus diversity and showed a high correlation (rSpearman = 0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics

DeeDeeExperiment: building an infrastructure for integrating and managing omics data analysis results in R/Bioconductor.

SUMMARY: Modern omics experiments now involve multiple conditions and complex designs, producing an increasingly large set of differential expression and functional enrichment analysis results. However, no standardized data structure exists to store and contextualize these results together with their metadata, leaving researchers with an unmanageable and potentially non-reproducible collection of results that are difficult to navigate and/or share. Here we introduce DeeDeeExperiment, a new S4 class for managing and storing omics data analysis results, implemented within the Bioconductor ecosystem, which promotes interoperability, reproducibility and good documentation. This class extends the widely used SingleCellExperiment object by introducing dedicated slots for Differential Expression (DEA) and Functional Enrichment Analysis (FEA) results, allowing users to organize, store, and retrieve information on multiple contrasts and associated metadata within a single data object, ultimately streamlining the management and interpretation of many omics datasets. AVAILABILITY AND IMPLEMENTATION: DeeDeeExperiment is available on Bioconductor under the MIT license (https://bioconductor.org/packages/DeeDeeExperiment), with its development version also available on Github (https://github.com/imbeimainz/DeeDeeExperiment).

Software

AWGE-ESPCA: An edge sparse PCA model based on adaptive noise elimination regularization and weighted gene network for Hermetia illucens genomic data analysis.

Hermetia illucens is an important insect resource. Studies have shown that exploring the effects of Cu2+-stressed on the growth and development of the Hermetia illucens genome holds significant scientific importance. There are three major challenges in the current studies of Hermetia illucens genomic data analysis: firstly, the lack of available genomic data which limits researchers in Hermetia illucens genomic data analysis. Secondly, to the best of our knowledge, there are no Artificial Intelligence (AI) feature selection models designed specifically for Hermetia illucens genome. Unlike human genomic data, noise in Hermetia illucens data is a more serious problem. Third, how to choose those genes located in the pathway enrichment region. Existing models assume that each gene probe has the same priori weight. However, researchers usually pay more attention to gene probes which are in the pathway enrichment region. Based on the above challenges, we initially construct experiments and establish a new Cu2+-stressed Hermetia illucens growth genome dataset. Subsequently, we propose AWGE-ESPCA: an edge Sparse PCA model based on adaptive noise elimination regularization and weighted gene network. The AWGE-ESPCA model innovatively proposes an adaptive noise elimination regularization method, effectively addressing the noise challenge in Hermetia illucens genomic data. We also integrate the known gene-pathway quantitative information into the Sparse PCA(SPCA) framework as a priori knowledge, which allows the model to filter out the gene probes in pathway-rich regions as much as possible. Ultimately, this study conducts five independent experiments and compared four latest Sparse PCA models as well as representative supervised and unsupervised baseline models to validate the model performance. The experimental results demonstrate the superior pathway and gene selection capabilities of the AWGE-ESPCA model. Ablation experiments validate the role of the adaptive regularizer and network weighting module. To summarize, this paper presents an innovative unsupervised model for Hermetia illucens genome analysis, which can effectively help researchers identify potential biomarkers. In addition, we also provide a working AWGE - ESPCA model code in the address: https://github.com/yhyresearcher/AWGE_ESPCA.

Animals

Augmented kurtosis-based projection pursuit: a novel, advanced machine learning approach for multi-omics data analysis and integration.

Due to the heterogeneity of multi-omics data, exacting their maximum information potential remains a challenge. Whereas some solutions have been offered, most cannot overcome the large linear dynamic range associated with such data, while others require large biological effect sizes to produce meaningful models. Here, we (i) perform a comprehensive benchmarking of multi-omics data analysis tools, and (ii) introduce kurtosis-based projection pursuit analysis, augmented with classification and regression trees (kPPA-CART) as a robust, easy-to-implement alternative. Using ground truth data, we demonstrate that kPPA-CART exhibits superiority in inferring biological significance from low-intensity (low-count) features and studies with small biological effect sizes. Applying it to experimental breast cancer data from The Cancer Genome Atlas, we identify novel genes that cluster the samples into subtypes that mimic the canonical PAM50 classes with notable improvements. Validating with external metastatic breast cancer data from the AURORA US consortium, kPPA-CART identifies genes that are associated with poor event-free survival and additional clustering associated with increased tumor mutational burden. Finally, we provide an R package and an online implementation of kPPA-CART.

Humans

A hierarchical, count-based model highlights challenges in scATAC-seq data analysis and points to opportunities to extract finer-resolution information.

BACKGROUND: Data from Single-cell Assay for Transposase Accessible Chromatin with Sequencing (scATAC-seq) is highly sparse. While current computational methods feature a range of transformation procedures to extract meaningful information, major challenges remain. RESULTS: Here, we discuss the major scATAC-seq data analysis challenges such as sequencing depth normalization and region-specific biases. We present a hierarchical count model that is motivated by the data generating process of scATAC-seq data. Our simulations show that current scATAC-seq data, while clearly containing physical single-cell resolution, are too sparse to infer true informational-level single-cell, single-region of chromatin accessibility states. CONCLUSIONS: While the broad utility of scATAC-seq at a cell type level is undeniable, describing it as fully resolving chromatin accessibility at single-cell resolution, particularly at individual locus level, may overstate the level of detail currently achievable. We conclude that chromatin accessibility profiling at true single-cell, single-region resolution is challenging with current data sensitivity, but that it may be achieved with promising developments in optimizing the efficiency of scATAC-seq assays.

Single-Cell Analysis

Implementing a training resource for large-scale genomic data analysis in the All of Us Researcher Workbench.

A lack of representation in genomic research and limited access to computational training create barriers for many researchers seeking to analyze large-scale genetic datasets. The All of Us Research Program provides an unprecedented opportunity to address these gaps by offering genomic data from a broad range of participants, but its impact depends on equipping researchers with the necessary skills to use it effectively. The All of Us Biomedical Researcher (BR) Scholars Program at Baylor College of Medicine aims to break down these barriers by providing early-career researchers with hands-on training in computational genomics through the All of Us Evenings with Genetics Research Program. The year-long program begins with the faculty summit, an in-person computational boot camp that introduces scholars to foundational skills for using the All of Us dataset via a cloud-based research environment. The genomics tutorials focus on genome-wide association studies (GWASs), utilizing Jupyter Notebooks and the Hail computing framework to provide an accessible and scalable approach to large-scale data analysis. Scholars engage in hands-on exercises covering data preparation, quality control, association testing, and result interpretation. By the end of the summit, participants will have successfully conducted a GWAS, visualized key findings, and gained confidence in computational resource management. This initiative expands access to genomic research by equipping early-career researchers from a variety of backgrounds with the tools and knowledge to analyze All of Us data. By lowering barriers to entry and promoting the study of representative populations, the program fosters innovation in precision medicine and advances equity in genomic research.

Humans

A simplified method of echocardiographic data analysis.

Rapid accurate analysis of echocardiographic data is accomplished using a sonic digitizer and programmable calculator. This method allows the echocardiographer to select technically optimal areas of the recording for analysis. The resolution of the measuring device is 0.1 mm. A hardcopy printout of both measurement and calculation is provided. Instead of expensive on-line computer, an inexpensive programmable calculator is used.

Computers

Analog processing of vestibular nystagmus for on-line cross- correlation data analysis.

An analog processing circuit is described which allow accurate measurement of the phase relationships between input angular acceleration and resulting eye velocity. Vestibular nystagmic data are processed via analog technics to yield slowphase eye velocity. The turntable velocity input is cross-correlated with the eye velocity output, using a Nicolet MED-80 minicomputer system. The resulting correlograms are further processed to obtain precise phase information. Test data analysis shows a system resolution within 1 degree. Data from human and animal subjects are portrayed.

Acceleration

Multivariate data analysis in empirical research. A look on the bright side.

The interpretive benefits of employing multivariate analysis methods on experimental data with more than one dependent variable are described heuristically and illustrated on a set of data from a simply designed experiment in physiological psychology. Multivariate analysis of variance (MANOVA) is performed on the 9 dependent variables contained in the sample data and on the four composites derived from a principal components analysis (PCA) of the variability of the nine. A linear discriminant analysis (LDA) is conducted following both MANOVA results, and 5 methods of determining the "important" dependent variables in the experimental-control group difference are presented and discussed in terms of the data at hand.

Analysis of Variance

HLA-D typing with lymphoblastoid cell lines. VII. A computer program for data analysis.

When lymphoblastoid cell lines (LCL) are substituted for peripheral blood lymphocytes from human typing cell donors in HLA-D typing experiments, a data analysis program must be designed to distinguish the effect of allo-reactivity from those peculiar to LCL, mainly the "autologous-stimulation" effect. The computer program described in this report was created specifically for such an analysis. The rationale for the design of this program is presented in the preceding report (see this issue).

Cell Line