PubMed HealthSearch

SEARCH · PubMed Health

Results for “Data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

UALCAN Mobile, an app for cancer proteogenomic data analysis.

Cancer is a complex disease affecting various organs and is a major cause of death worldwide. During cancer initiation, disease progression, and tumor metastasis, various genomic and proteomic alterations are observed. Recent technological advances have led to the generation of large amounts of molecular data, including genomics and transcriptomics. These large-scale datasets can be utilized to analyze and identify sub-class-specific cancer biomarkers and targets. However, there is a need for the development of user-friendly tools for large-scale data analysis, disseminating the analyzed data in a visualizable format to cancer researchers with no programming skills. We developed UALCAN, a comprehensive platform that allows users to integrate disparate data to better understand the genes, proteins, and pathways perturbed in cancer and make discoveries of potential biomarkers and targets. In the current study, we describe the development of the UALCAN Mobile application (app) that will provide cancer transcriptomic data obtained from The Cancer Genome Atlas (TCGA) project to evaluate protein-coding gene expression based on various stratifications, including stage, grade, race, gender, and molecular-subtypes across over 30 types of cancers. In addition, the UALCAN mobile provides data analysis options for epigenetic changes due to DNA promoter methylation and Clinical Proteomic Tumor Analysis Consortium (CPTAC) cancer proteomic data. The app provides access to large cancer molecular datasets on the go. To find changes in the expression of causative genes and proteins and to identify biomarkers and therapeutic targets, UALCAN mobile app will be extremely valuable. The "UALCAN Mobile" app is free to use and can be downloaded from both the iOS/Apple and the Android Play Store and has been downloaded over 100 times in each of iOS and android app stores.

app

Research in physical medicine and rehabilitation. V. Data entry and early exploratory data analysis.

The process of data entry and initial analysis to locate data errors is described. Basic terms are defined and a simple method of entering data by using word processing software is illustrated. Data checking is done by using visual check of the raw data. Statistical programs are then used to locate possible data errors by finding data points (outliers) that are very different from the average. Special graphic output of statistical programs, scatterplots and box and whisker plots can be used to further locate questionable data. Examples of data entry forms and annotated step by step data cleaning with the use of inexpensive programs for personal computers are presented.

Computers

Uncovering psychiatric test information with graphical techniques of Exploratory Data Analysis.

This article illustrates how Exploratory Data Analysis (EDA) can complement conventional statistical methods in evaluating psychiatric tests. Using one recent EDA computer program, we evaluated the ability of repeated psychiatric screening tests (the General Health Questionnaire [GHQ]) to predict medical and psychiatric service use in a Health Maintenance Organization (HMO), the Harvard Community Health Plan (HCHP). Using a stratified random sample of 244 new HCHP enrollees and viewing three-dimensional graphs of their data from multiple perspectives, we found two subpopulations: low GHQ scorers, for whom the tests did not predict service use; and high scorers, for whom they did. Surprisingly, improving scores forecast increased use and chronically high scores predicted diminished use. Using another stratified random sample of 213 new HCHP enrollees, and with scatterplot matrices from another interactive computer program, we found that high and unchanging GHQ scores forecast HMO dropout. We examine possible interpretations--for example, that chronically distressed patients may become immobilized, diminish service use, and ultimately leave the HMO. We also explain how EDA methods may help uncover elusive results in other data (e.g., mental health outcomes).

Adult

Research in physical medicine and rehabilitation. VIII. Preliminary data analysis.

This paper describes important aspects of preliminary data analysis to be taken after data are checked for clerical entry errors and before the primary statistical analysis is performed. These include description and graphic display of each variable, recoding categorical data, transforming continuous data into another continuous variable and recoding continuous to categorical data. Missing values and outlying data points are identified and several techniques are recommended to minimize mistakes in variable recoding. Related variables measured with different units may be combined by using the z transformation and converted back to one of the original units for ease of interpretation. Finally, both categorical and continuous variables are checked for reliability by using kappa or the intraclass R.

Data Collection

BioNeuralNet: a graph neural network based Multi-Omics network data analysis tool.

SUMMARY: Multi-omics data offer unprecedented insights into complex biological systems, yet their high dimensionality, sparsity, and intricate interactions pose significant analytical challenges. Network-based approaches have advanced multi-omics research by effectively capturing biologically relevant relationships among molecular features (e.g., genes, proteins, metabolites). While these methods are powerful for representing molecular interactions, there remains a need for tools specifically designed to effectively utilize these network representations across diverse downstream analyses. To fulfill this need, we introduce BioNeuralNet, a flexible and modular Python framework tailored for end-to-end network-based multi-omics data analysis. BioNeuralNet leverages Graph Neural Networks (GNNs) to learn biologically meaningful low-dimensional representations from multi-omics networks, converting these complex molecular networks into versatile embeddings. BioNeuralNet supports all major stages of multi-omics network analysis, including several network construction techniques, generation of low-dimensional representations, and a broad range of downstream analytical tasks. Its extensive utilities, including diverse GNN architectures, and compatibility with established Python packages (e.g., scikit-learn, PyTorch, NetworkX), enhance usability and facilitate quick adoption. BioNeuralNet is an open-source, user-friendly, and extensively documented framework designed to support flexible and reproducible multi-omics network analysis in precision medicine. AVAILABILITY AND IMPLEMENTATION: The BioNeuralNet library is available via The Python Package Index (PyPI). Source code, documentation, tutorials, and workflows are hosted at https://bioneuralnet.readthedocs.io. Code archived at https://doi.org/10.5281/zenodo.17503083.

Graph Neural Networks

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score = 0.67-0.90) in genus diversity and showed a high correlation (rSpearman = 0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics

Computer-assisted diagnosis by a model-free system of direct data analysis.

The basis of the method of data analysis presented is, in the case of any diagnostic test, the automatic compilation of separate frequency distributions for each diagnostic classification. The distinction of different test results for different diseases (the correlation for which the tests are used) can thus be quantitatively monitored. This offers opportunities for more specific control of the accuracy of the data base. Measurements of relative frequencies obtained from the frequency distributions of individuals with and without a given disease can serve as a quantitative handle for the selection of the combination of tests, and for adjustments of individual parameters, which will maximize the discrimination. The usual cutoffs are not used. A data-processing system can serve for the direct incorporation of patient chart data (including test results), and for the automation of the analysis described, with pattern recognition or cluster-seeking techniques. The ability of this system of analysis to minimize some of the problems associated with methods utilizing mathematical models is discussed.

Diagnosis, Computer-Assisted

Toward a computer assisted analysis of NOESY spectra: a multivariate data analysis of an RNA NOESY spectrum.

A multivariate data-representation of a portion of the H-NOESY spectrum of an RNA octamer duplex was used to explore the possibility of using Principal Component Analysis and Partial Least Squares Discrimination for pattern recognition. In this case, it is found that the methods can: (i) distinguish slices containing signal from those containing only noise, (ii) locate slices containing overlapping signals, and (iii) in some cases to segregate slices with unique aspects such as those from terminal nucleotides, overlapping signals, purine-H8, pyrimidine-H6 and adenine-H2 containing slices. These properties can easily be included in a scheme to automate spectral analysis. The formulation described here does not distinguish patterns needed to automate sequential assignment of resonances in NOESY spectra of RNA.

Base Sequence

DeeDeeExperiment: building an infrastructure for integrating and managing omics data analysis results in R/Bioconductor.

SUMMARY: Modern omics experiments now involve multiple conditions and complex designs, producing an increasingly large set of differential expression and functional enrichment analysis results. However, no standardized data structure exists to store and contextualize these results together with their metadata, leaving researchers with an unmanageable and potentially non-reproducible collection of results that are difficult to navigate and/or share. Here we introduce DeeDeeExperiment, a new S4 class for managing and storing omics data analysis results, implemented within the Bioconductor ecosystem, which promotes interoperability, reproducibility and good documentation. This class extends the widely used SingleCellExperiment object by introducing dedicated slots for Differential Expression (DEA) and Functional Enrichment Analysis (FEA) results, allowing users to organize, store, and retrieve information on multiple contrasts and associated metadata within a single data object, ultimately streamlining the management and interpretation of many omics datasets. AVAILABILITY AND IMPLEMENTATION: DeeDeeExperiment is available on Bioconductor under the MIT license (https://bioconductor.org/packages/DeeDeeExperiment), with its development version also available on Github (https://github.com/imbeimainz/DeeDeeExperiment).

Software

Planning controlled clinical trials on the basis of descriptive data analysis.

In controlled clinical trials the problem of multiplicity of desired inferential statements finds attention at an increasing rate. In this paper the previously proposed concept of Descriptive Data Analysis (DDA), situated between Confirmatory and Exploratory Data Analysis, is applied to the planning aspects of controlled trials for which the problem of multiplicity exists. The (non-Bayesian) DDA planning concept should provide the investigator with tools to draw final conclusions from data of several variables possibly observed at several time points in possibly several groups of subjects by combining his pre-trial medical experience with descriptive inferential statements (confidence intervals and test results) at nominal significance levels. DDA also provides for confirmatory statements concerning individual null hypotheses and partially global hypotheses.

Clinical Trials as Topic

AWGE-ESPCA: An edge sparse PCA model based on adaptive noise elimination regularization and weighted gene network for Hermetia illucens genomic data analysis.

Hermetia illucens is an important insect resource. Studies have shown that exploring the effects of Cu2+-stressed on the growth and development of the Hermetia illucens genome holds significant scientific importance. There are three major challenges in the current studies of Hermetia illucens genomic data analysis: firstly, the lack of available genomic data which limits researchers in Hermetia illucens genomic data analysis. Secondly, to the best of our knowledge, there are no Artificial Intelligence (AI) feature selection models designed specifically for Hermetia illucens genome. Unlike human genomic data, noise in Hermetia illucens data is a more serious problem. Third, how to choose those genes located in the pathway enrichment region. Existing models assume that each gene probe has the same priori weight. However, researchers usually pay more attention to gene probes which are in the pathway enrichment region. Based on the above challenges, we initially construct experiments and establish a new Cu2+-stressed Hermetia illucens growth genome dataset. Subsequently, we propose AWGE-ESPCA: an edge Sparse PCA model based on adaptive noise elimination regularization and weighted gene network. The AWGE-ESPCA model innovatively proposes an adaptive noise elimination regularization method, effectively addressing the noise challenge in Hermetia illucens genomic data. We also integrate the known gene-pathway quantitative information into the Sparse PCA(SPCA) framework as a priori knowledge, which allows the model to filter out the gene probes in pathway-rich regions as much as possible. Ultimately, this study conducts five independent experiments and compared four latest Sparse PCA models as well as representative supervised and unsupervised baseline models to validate the model performance. The experimental results demonstrate the superior pathway and gene selection capabilities of the AWGE-ESPCA model. Ablation experiments validate the role of the adaptive regularizer and network weighting module. To summarize, this paper presents an innovative unsupervised model for Hermetia illucens genome analysis, which can effectively help researchers identify potential biomarkers. In addition, we also provide a working AWGE - ESPCA model code in the address: https://github.com/yhyresearcher/AWGE_ESPCA.

Animals

Database and search techniques for two-dimensional gel protein data: a comparison of paradigms for exploratory data analysis and prospects for biological modeling.

Two-dimensional (2-D) polyacrylamide gel electrophoresis can detect thousands of polypeptides, separating them by apparent molecular weight (Mr) and isoelectric point (pI). Thus it provides a more realistic and global view of cellular genetic expression than any other technique. This technique has been useful for finding sets of key proteins of biological significance. However, a typical experiment with more than a few gels often results in an unwiedly data management problem. In this paper, the GELLAB-II system is discussed with respect to how data reduction and exploratory data analysis can be aided by computer data management and statistical search techniques. By encoding the gel patterns in a "three-dimensional" (3-D) database, an exploratory data analysis can be carried out in an environment that might be called a "spread sheet for 2-D gel protein data". From such databases, complex parametric network models of protein expression during events such as differentiation might be constructed. For this, 2-D gel databases must be able to include data from other domains external to the gel itself. Because of the increasing complexity of such databases, new tools are required to help manage this complexity. Two such tools, object-oriented databases and expert-system rule-based analysis, are discussed in this context. Comparisons are made between GELLAB and other 2-D gel database analysis systems to illustrate some of the analysis paradigms common to these systems and where this technology may be heading.

Algorithms

Area normalization of the renal region of interest in radionuclide renography data analysis: a misconception.

Relative renal function is estimated by comparing the area under the second segment of the curve from the renal region of interest in a renographic study. We have examined the problems arising out of area normalization of the renal region of interest in the data analysis for relative renal function evaluation. Error analysis by computer simulation proves that this method of data analysis is highly misleading and erroneous.

Humans

Empirical considerations in orthopaedic research design and data analysis. Part II: The application of data analytic techniques.

To assure that a hypothesis is tested as rigorously as possible, the proper statistical method must be used to analyze the data. But without a strong background in statistics, it may be difficult to determine the efficacy of the data analytic technique used in the study. This paper describes several widely used data analytic techniques and offers examples of their proper application in orthopaedic research design.

Data Interpretation, Statistical

Augmented kurtosis-based projection pursuit: a novel, advanced machine learning approach for multi-omics data analysis and integration.

Due to the heterogeneity of multi-omics data, exacting their maximum information potential remains a challenge. Whereas some solutions have been offered, most cannot overcome the large linear dynamic range associated with such data, while others require large biological effect sizes to produce meaningful models. Here, we (i) perform a comprehensive benchmarking of multi-omics data analysis tools, and (ii) introduce kurtosis-based projection pursuit analysis, augmented with classification and regression trees (kPPA-CART) as a robust, easy-to-implement alternative. Using ground truth data, we demonstrate that kPPA-CART exhibits superiority in inferring biological significance from low-intensity (low-count) features and studies with small biological effect sizes. Applying it to experimental breast cancer data from The Cancer Genome Atlas, we identify novel genes that cluster the samples into subtypes that mimic the canonical PAM50 classes with notable improvements. Validating with external metastatic breast cancer data from the AURORA US consortium, kPPA-CART identifies genes that are associated with poor event-free survival and additional clustering associated with increased tumor mutational burden. Finally, we provide an R package and an online implementation of kPPA-CART.

Humans

A hierarchical, count-based model highlights challenges in scATAC-seq data analysis and points to opportunities to extract finer-resolution information.

BACKGROUND: Data from Single-cell Assay for Transposase Accessible Chromatin with Sequencing (scATAC-seq) is highly sparse. While current computational methods feature a range of transformation procedures to extract meaningful information, major challenges remain. RESULTS: Here, we discuss the major scATAC-seq data analysis challenges such as sequencing depth normalization and region-specific biases. We present a hierarchical count model that is motivated by the data generating process of scATAC-seq data. Our simulations show that current scATAC-seq data, while clearly containing physical single-cell resolution, are too sparse to infer true informational-level single-cell, single-region of chromatin accessibility states. CONCLUSIONS: While the broad utility of scATAC-seq at a cell type level is undeniable, describing it as fully resolving chromatin accessibility at single-cell resolution, particularly at individual locus level, may overstate the level of detail currently achievable. We conclude that chromatin accessibility profiling at true single-cell, single-region resolution is challenging with current data sensitivity, but that it may be achieved with promising developments in optimizing the efficiency of scATAC-seq assays.

Single-Cell Analysis

[The effect of smoking habit on aortic pulse wave velocity using a new method for data analysis].

We measured aortic pulse wave velocity (PWV) in 168 male adult cases of various arteriosclerotic diseases. In order to evaluate the effects of age, smoking habits, alcohol intake, and blood pressure, we applied the least median of squares (LMS) regression which was considered to be very useful for data analysis. The results showed that PWV level increased with age. Furthermore smoking was associated with increasing PWV level and this effect was also related to age. We concluded that the PWV was valuable as an index of arteriosclerosis, and instead of the classical least squares method, LMS regression was very useful for analysis of medical data.

Adult

Implementing a training resource for large-scale genomic data analysis in the All of Us Researcher Workbench.

A lack of representation in genomic research and limited access to computational training create barriers for many researchers seeking to analyze large-scale genetic datasets. The All of Us Research Program provides an unprecedented opportunity to address these gaps by offering genomic data from a broad range of participants, but its impact depends on equipping researchers with the necessary skills to use it effectively. The All of Us Biomedical Researcher (BR) Scholars Program at Baylor College of Medicine aims to break down these barriers by providing early-career researchers with hands-on training in computational genomics through the All of Us Evenings with Genetics Research Program. The year-long program begins with the faculty summit, an in-person computational boot camp that introduces scholars to foundational skills for using the All of Us dataset via a cloud-based research environment. The genomics tutorials focus on genome-wide association studies (GWASs), utilizing Jupyter Notebooks and the Hail computing framework to provide an accessible and scalable approach to large-scale data analysis. Scholars engage in hands-on exercises covering data preparation, quality control, association testing, and result interpretation. By the end of the summit, participants will have successfully conducted a GWAS, visualized key findings, and gained confidence in computational resource management. This initiative expands access to genomic research by equipping early-career researchers from a variety of backgrounds with the tools and knowledge to analyze All of Us data. By lowering barriers to entry and promoting the study of representative populations, the program fosters innovation in precision medicine and advances equity in genomic research.

Humans