PubMed HealthSearch

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Uncovering psychiatric test information with graphical techniques of Exploratory Data Analysis.

This article illustrates how Exploratory Data Analysis (EDA) can complement conventional statistical methods in evaluating psychiatric tests. Using one recent EDA computer program, we evaluated the ability of repeated psychiatric screening tests (the General Health Questionnaire [GHQ]) to predict medical and psychiatric service use in a Health Maintenance Organization (HMO), the Harvard Community Health Plan (HCHP). Using a stratified random sample of 244 new HCHP enrollees and viewing three-dimensional graphs of their data from multiple perspectives, we found two subpopulations: low GHQ scorers, for whom the tests did not predict service use; and high scorers, for whom they did. Surprisingly, improving scores forecast increased use and chronically high scores predicted diminished use. Using another stratified random sample of 213 new HCHP enrollees, and with scatterplot matrices from another interactive computer program, we found that high and unchanging GHQ scores forecast HMO dropout. We examine possible interpretations--for example, that chronically distressed patients may become immobilized, diminish service use, and ultimately leave the HMO. We also explain how EDA methods may help uncover elusive results in other data (e.g., mental health outcomes).

Adult

Research in physical medicine and rehabilitation. VIII. Preliminary data analysis.

This paper describes important aspects of preliminary data analysis to be taken after data are checked for clerical entry errors and before the primary statistical analysis is performed. These include description and graphic display of each variable, recoding categorical data, transforming continuous data into another continuous variable and recoding continuous to categorical data. Missing values and outlying data points are identified and several techniques are recommended to minimize mistakes in variable recoding. Related variables measured with different units may be combined by using the z transformation and converted back to one of the original units for ease of interpretation. Finally, both categorical and continuous variables are checked for reliability by using kappa or the intraclass R.

Data Collection

BioNeuralNet: a graph neural network based Multi-Omics network data analysis tool.

SUMMARY: Multi-omics data offer unprecedented insights into complex biological systems, yet their high dimensionality, sparsity, and intricate interactions pose significant analytical challenges. Network-based approaches have advanced multi-omics research by effectively capturing biologically relevant relationships among molecular features (e.g., genes, proteins, metabolites). While these methods are powerful for representing molecular interactions, there remains a need for tools specifically designed to effectively utilize these network representations across diverse downstream analyses. To fulfill this need, we introduce BioNeuralNet, a flexible and modular Python framework tailored for end-to-end network-based multi-omics data analysis. BioNeuralNet leverages Graph Neural Networks (GNNs) to learn biologically meaningful low-dimensional representations from multi-omics networks, converting these complex molecular networks into versatile embeddings. BioNeuralNet supports all major stages of multi-omics network analysis, including several network construction techniques, generation of low-dimensional representations, and a broad range of downstream analytical tasks. Its extensive utilities, including diverse GNN architectures, and compatibility with established Python packages (e.g., scikit-learn, PyTorch, NetworkX), enhance usability and facilitate quick adoption. BioNeuralNet is an open-source, user-friendly, and extensively documented framework designed to support flexible and reproducible multi-omics network analysis in precision medicine. AVAILABILITY AND IMPLEMENTATION: The BioNeuralNet library is available via The Python Package Index (PyPI). Source code, documentation, tutorials, and workflows are hosted at https://bioneuralnet.readthedocs.io. Code archived at https://doi.org/10.5281/zenodo.17503083.

Graph Neural Networks

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score = 0.67-0.90) in genus diversity and showed a high correlation (rSpearman = 0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics

Computer-assisted diagnosis by a model-free system of direct data analysis.

The basis of the method of data analysis presented is, in the case of any diagnostic test, the automatic compilation of separate frequency distributions for each diagnostic classification. The distinction of different test results for different diseases (the correlation for which the tests are used) can thus be quantitatively monitored. This offers opportunities for more specific control of the accuracy of the data base. Measurements of relative frequencies obtained from the frequency distributions of individuals with and without a given disease can serve as a quantitative handle for the selection of the combination of tests, and for adjustments of individual parameters, which will maximize the discrimination. The usual cutoffs are not used. A data-processing system can serve for the direct incorporation of patient chart data (including test results), and for the automation of the analysis described, with pattern recognition or cluster-seeking techniques. The ability of this system of analysis to minimize some of the problems associated with methods utilizing mathematical models is discussed.

Diagnosis, Computer-Assisted

DeeDeeExperiment: building an infrastructure for integrating and managing omics data analysis results in R/Bioconductor.

SUMMARY: Modern omics experiments now involve multiple conditions and complex designs, producing an increasingly large set of differential expression and functional enrichment analysis results. However, no standardized data structure exists to store and contextualize these results together with their metadata, leaving researchers with an unmanageable and potentially non-reproducible collection of results that are difficult to navigate and/or share. Here we introduce DeeDeeExperiment, a new S4 class for managing and storing omics data analysis results, implemented within the Bioconductor ecosystem, which promotes interoperability, reproducibility and good documentation. This class extends the widely used SingleCellExperiment object by introducing dedicated slots for Differential Expression (DEA) and Functional Enrichment Analysis (FEA) results, allowing users to organize, store, and retrieve information on multiple contrasts and associated metadata within a single data object, ultimately streamlining the management and interpretation of many omics datasets. AVAILABILITY AND IMPLEMENTATION: DeeDeeExperiment is available on Bioconductor under the MIT license (https://bioconductor.org/packages/DeeDeeExperiment), with its development version also available on Github (https://github.com/imbeimainz/DeeDeeExperiment).

Software

Planning controlled clinical trials on the basis of descriptive data analysis.

In controlled clinical trials the problem of multiplicity of desired inferential statements finds attention at an increasing rate. In this paper the previously proposed concept of Descriptive Data Analysis (DDA), situated between Confirmatory and Exploratory Data Analysis, is applied to the planning aspects of controlled trials for which the problem of multiplicity exists. The (non-Bayesian) DDA planning concept should provide the investigator with tools to draw final conclusions from data of several variables possibly observed at several time points in possibly several groups of subjects by combining his pre-trial medical experience with descriptive inferential statements (confidence intervals and test results) at nominal significance levels. DDA also provides for confirmatory statements concerning individual null hypotheses and partially global hypotheses.

Clinical Trials as Topic

AWGE-ESPCA: An edge sparse PCA model based on adaptive noise elimination regularization and weighted gene network for Hermetia illucens genomic data analysis.

Hermetia illucens is an important insect resource. Studies have shown that exploring the effects of Cu2+-stressed on the growth and development of the Hermetia illucens genome holds significant scientific importance. There are three major challenges in the current studies of Hermetia illucens genomic data analysis: firstly, the lack of available genomic data which limits researchers in Hermetia illucens genomic data analysis. Secondly, to the best of our knowledge, there are no Artificial Intelligence (AI) feature selection models designed specifically for Hermetia illucens genome. Unlike human genomic data, noise in Hermetia illucens data is a more serious problem. Third, how to choose those genes located in the pathway enrichment region. Existing models assume that each gene probe has the same priori weight. However, researchers usually pay more attention to gene probes which are in the pathway enrichment region. Based on the above challenges, we initially construct experiments and establish a new Cu2+-stressed Hermetia illucens growth genome dataset. Subsequently, we propose AWGE-ESPCA: an edge Sparse PCA model based on adaptive noise elimination regularization and weighted gene network. The AWGE-ESPCA model innovatively proposes an adaptive noise elimination regularization method, effectively addressing the noise challenge in Hermetia illucens genomic data. We also integrate the known gene-pathway quantitative information into the Sparse PCA(SPCA) framework as a priori knowledge, which allows the model to filter out the gene probes in pathway-rich regions as much as possible. Ultimately, this study conducts five independent experiments and compared four latest Sparse PCA models as well as representative supervised and unsupervised baseline models to validate the model performance. The experimental results demonstrate the superior pathway and gene selection capabilities of the AWGE-ESPCA model. Ablation experiments validate the role of the adaptive regularizer and network weighting module. To summarize, this paper presents an innovative unsupervised model for Hermetia illucens genome analysis, which can effectively help researchers identify potential biomarkers. In addition, we also provide a working AWGE - ESPCA model code in the address: https://github.com/yhyresearcher/AWGE_ESPCA.

Animals

Area normalization of the renal region of interest in radionuclide renography data analysis: a misconception.

Relative renal function is estimated by comparing the area under the second segment of the curve from the renal region of interest in a renographic study. We have examined the problems arising out of area normalization of the renal region of interest in the data analysis for relative renal function evaluation. Error analysis by computer simulation proves that this method of data analysis is highly misleading and erroneous.

Humans

Empirical considerations in orthopaedic research design and data analysis. Part II: The application of data analytic techniques.

To assure that a hypothesis is tested as rigorously as possible, the proper statistical method must be used to analyze the data. But without a strong background in statistics, it may be difficult to determine the efficacy of the data analytic technique used in the study. This paper describes several widely used data analytic techniques and offers examples of their proper application in orthopaedic research design.

Data Interpretation, Statistical

Augmented kurtosis-based projection pursuit: a novel, advanced machine learning approach for multi-omics data analysis and integration.

Due to the heterogeneity of multi-omics data, exacting their maximum information potential remains a challenge. Whereas some solutions have been offered, most cannot overcome the large linear dynamic range associated with such data, while others require large biological effect sizes to produce meaningful models. Here, we (i) perform a comprehensive benchmarking of multi-omics data analysis tools, and (ii) introduce kurtosis-based projection pursuit analysis, augmented with classification and regression trees (kPPA-CART) as a robust, easy-to-implement alternative. Using ground truth data, we demonstrate that kPPA-CART exhibits superiority in inferring biological significance from low-intensity (low-count) features and studies with small biological effect sizes. Applying it to experimental breast cancer data from The Cancer Genome Atlas, we identify novel genes that cluster the samples into subtypes that mimic the canonical PAM50 classes with notable improvements. Validating with external metastatic breast cancer data from the AURORA US consortium, kPPA-CART identifies genes that are associated with poor event-free survival and additional clustering associated with increased tumor mutational burden. Finally, we provide an R package and an online implementation of kPPA-CART.

Humans

A hierarchical, count-based model highlights challenges in scATAC-seq data analysis and points to opportunities to extract finer-resolution information.

BACKGROUND: Data from Single-cell Assay for Transposase Accessible Chromatin with Sequencing (scATAC-seq) is highly sparse. While current computational methods feature a range of transformation procedures to extract meaningful information, major challenges remain. RESULTS: Here, we discuss the major scATAC-seq data analysis challenges such as sequencing depth normalization and region-specific biases. We present a hierarchical count model that is motivated by the data generating process of scATAC-seq data. Our simulations show that current scATAC-seq data, while clearly containing physical single-cell resolution, are too sparse to infer true informational-level single-cell, single-region of chromatin accessibility states. CONCLUSIONS: While the broad utility of scATAC-seq at a cell type level is undeniable, describing it as fully resolving chromatin accessibility at single-cell resolution, particularly at individual locus level, may overstate the level of detail currently achievable. We conclude that chromatin accessibility profiling at true single-cell, single-region resolution is challenging with current data sensitivity, but that it may be achieved with promising developments in optimizing the efficiency of scATAC-seq assays.

Single-Cell Analysis

Implementing a training resource for large-scale genomic data analysis in the All of Us Researcher Workbench.

A lack of representation in genomic research and limited access to computational training create barriers for many researchers seeking to analyze large-scale genetic datasets. The All of Us Research Program provides an unprecedented opportunity to address these gaps by offering genomic data from a broad range of participants, but its impact depends on equipping researchers with the necessary skills to use it effectively. The All of Us Biomedical Researcher (BR) Scholars Program at Baylor College of Medicine aims to break down these barriers by providing early-career researchers with hands-on training in computational genomics through the All of Us Evenings with Genetics Research Program. The year-long program begins with the faculty summit, an in-person computational boot camp that introduces scholars to foundational skills for using the All of Us dataset via a cloud-based research environment. The genomics tutorials focus on genome-wide association studies (GWASs), utilizing Jupyter Notebooks and the Hail computing framework to provide an accessible and scalable approach to large-scale data analysis. Scholars engage in hands-on exercises covering data preparation, quality control, association testing, and result interpretation. By the end of the summit, participants will have successfully conducted a GWAS, visualized key findings, and gained confidence in computational resource management. This initiative expands access to genomic research by equipping early-career researchers from a variety of backgrounds with the tools and knowledge to analyze All of Us data. By lowering barriers to entry and promoting the study of representative populations, the program fosters innovation in precision medicine and advances equity in genomic research.

Humans

Proficiency of the Tradescantia-micronucleus image analysis system for scoring micronucleus frequencies and data analysis.

The Tradescantia-micronucleus (Trad-MCN) bioassay is an efficient short-term test for genotoxicity of pollutants. In order to increase the efficiency and to standardize the micronucleus (MCN) scoring process, an automated scoring system was developed using the principle of image analysis in computer science. This assemblage is called the Tradescantia-micronucleus image analysis (Trad-MCNIA) system. The MCN frequencies scored by this system were compared with those scored by human observation for its proficiency. A set of low MCN frequency (around 5 MCN/100 tetrads) slides prepared from a control group, a set of medium MCN frequency (around 20 MCN/100 tetrads) slides prepared from sodium azide treated plant cuttings and a set of high MCN frequency (around 50 MCN/100 tetrads) slides prepared from X-ray treated materials were used for this study. In the low MCN frequency slides, the Trad-MCNIA system scored about the same value as human observation. In the medium and high frequency slides, MCN frequencies scored by the system were lower than those scored by human observers. This discrepancy was corrected by increasing the power of the objective of the microscope in the system. The MCN frequencies scored by the system attained 90% congruity with those scored by human observers after the correction. The scoring speed of the system was about 3.5 times as fast as that by human observers, and the data could be statistically analyzed immediately after the data scores were recorded. Further improvements can be made by upgrading the video camera and the computer speed.

Azides

A simplified method of echocardiographic data analysis.

Rapid accurate analysis of echocardiographic data is accomplished using a sonic digitizer and programmable calculator. This method allows the echocardiographer to select technically optimal areas of the recording for analysis. The resolution of the measuring device is 0.1 mm. A hardcopy printout of both measurement and calculation is provided. Instead of expensive on-line computer, an inexpensive programmable calculator is used.

Computers

A consultation system constructor for medical data analysis.

MAD is a system that helps an expert data analyst in a specific application domain (like epidemiology or image analysis) to build reasoning models aimed at fulfilling specific tasks. These models may be subsequently used to guide doctors in the analysis of a set of data referring to a specific ground domain. Expert knowledge is represented at various levels: a general description of an application domain and various models that formalize the reasoning followed to perform specific tasks within a defined application domain. Reasoning models are represented as rules of propositional calculus, and a meta-knowledge permits to support knowledge acquisition. During the consultation, different external programs may be run when needed, without the doctor having to learn how to use them. MAD is written in Golden Common LISP and may be linked to any external software for data analysis, provided it runs under MS-DOS and does not require more than 192 Kb. Examples of application of the system to epidemiology and image analysis are given.

Computer Simulation

Constrained and restrained refinement in EXAFS data analysis with curved wave theory.

This paper describes methods of constrained and restrained refinement of EXAFS data which provide a means of substantially reducing the number of independent parameters compared to conventional least-squares methods commonly used. Constrained refinement allows a major reduction in the number of free parameters for a refinement of a structural model. In restrained refinement, additional structural information from well-characterized small molecules is used to provide additional observations in the data analysis. Even though these methods are of general application to the majority of complex systems, they are particularly valuable for biological molecules. The methods are of major advantage for ligands where significant multiple scattering is present, e.g., histidine, tyrosine, CO, CN, etc. The bases of these methods are described, and applications to some complex chemical and biological systems are given.

Fetal Hemoglobin

Analog processing of vestibular nystagmus for on-line cross- correlation data analysis.

An analog processing circuit is described which allow accurate measurement of the phase relationships between input angular acceleration and resulting eye velocity. Vestibular nystagmic data are processed via analog technics to yield slowphase eye velocity. The turntable velocity input is cross-correlated with the eye velocity output, using a Nicolet MED-80 minicomputer system. The resulting correlograms are further processed to obtain precise phase information. Test data analysis shows a system resolution within 1 degree. Data from human and animal subjects are portrayed.

Acceleration