PubMed HealthSearch

SEARCH · PubMed Health

Results for “Data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Empirical considerations in orthopaedic research design and data analysis. Part II: The application of data analytic techniques.

To assure that a hypothesis is tested as rigorously as possible, the proper statistical method must be used to analyze the data. But without a strong background in statistics, it may be difficult to determine the efficacy of the data analytic technique used in the study. This paper describes several widely used data analytic techniques and offers examples of their proper application in orthopaedic research design.

Data Interpretation, Statistical

A categorical data analysis of contacts with the Family Health Clinic, Calabar, Nigeria.

The relationships of population, environmental and accessibility variables to registration and attendance by mothers of children under 6 at the Family Health Clinic in Calabar, Nigeria are investigated. The technique used to analyze the data collected is categorical data analysis which proceeds in two stages, variable selection to reduce the variable set and fitting a log-linear model to the reduced set. Details of the statistical procedures used are provided to indicate how categorical data analysis can be used as a valuable tool of analysis in medical geographical studies that employ count or frequency data. It was found that younger mothers and Ibibio women registered more often at the clinic than did their counterparts. However, if the relatively sparse data on fathers is accepted, the association between age and registration is found to be spurious and a model can be substituted which shows younger fathers and fathers who spoke a non-Efik/Ibibio language to be associated with higher clinic registration of mothers. It was further found that for registered mothers the probability of a clinic visit was decreased by mother's age, increased by distance given no travel cost, unaffected by distance given some travel cost, increased by travel cost given a short distance to the clinic and decreased by travel cost given a longer distance from the clinic. These results are discussed in relation to population characteristics such as socio-economic status, clinic procedures such as health worker activities, transportation availability in Calabar, the spatial ecology of the city and local environmental conditions.

Adult

Augmented kurtosis-based projection pursuit: a novel, advanced machine learning approach for multi-omics data analysis and integration.

Due to the heterogeneity of multi-omics data, exacting their maximum information potential remains a challenge. Whereas some solutions have been offered, most cannot overcome the large linear dynamic range associated with such data, while others require large biological effect sizes to produce meaningful models. Here, we (i) perform a comprehensive benchmarking of multi-omics data analysis tools, and (ii) introduce kurtosis-based projection pursuit analysis, augmented with classification and regression trees (kPPA-CART) as a robust, easy-to-implement alternative. Using ground truth data, we demonstrate that kPPA-CART exhibits superiority in inferring biological significance from low-intensity (low-count) features and studies with small biological effect sizes. Applying it to experimental breast cancer data from The Cancer Genome Atlas, we identify novel genes that cluster the samples into subtypes that mimic the canonical PAM50 classes with notable improvements. Validating with external metastatic breast cancer data from the AURORA US consortium, kPPA-CART identifies genes that are associated with poor event-free survival and additional clustering associated with increased tumor mutational burden. Finally, we provide an R package and an online implementation of kPPA-CART.

Humans

Novel data analysis for synchronised spontaneous neuromagnetic activity.

A novel approach to neuromagnetic data analysis is presented. This technique is aimed at studying synchronised spontaneous activity (SSA) and has been used to resolve two different signals from one single evoked response, providing evidence for two possibly distinct sources. The data presented are consistent with a model that permits the generators of spontaneous activity to be synchronised by sensory stimuli.

Brain

A hierarchical, count-based model highlights challenges in scATAC-seq data analysis and points to opportunities to extract finer-resolution information.

BACKGROUND: Data from Single-cell Assay for Transposase Accessible Chromatin with Sequencing (scATAC-seq) is highly sparse. While current computational methods feature a range of transformation procedures to extract meaningful information, major challenges remain. RESULTS: Here, we discuss the major scATAC-seq data analysis challenges such as sequencing depth normalization and region-specific biases. We present a hierarchical count model that is motivated by the data generating process of scATAC-seq data. Our simulations show that current scATAC-seq data, while clearly containing physical single-cell resolution, are too sparse to infer true informational-level single-cell, single-region of chromatin accessibility states. CONCLUSIONS: While the broad utility of scATAC-seq at a cell type level is undeniable, describing it as fully resolving chromatin accessibility at single-cell resolution, particularly at individual locus level, may overstate the level of detail currently achievable. We conclude that chromatin accessibility profiling at true single-cell, single-region resolution is challenging with current data sensitivity, but that it may be achieved with promising developments in optimizing the efficiency of scATAC-seq assays.

Single-Cell Analysis

[Computer-assisted data analysis in a pediatric intensive care unit].

Computer assisted real time data analysis introduces a reasonable method of judgment into patient monitoring systems. From fast changing vital parameters discrete heart and respiration rate samples are immediately evaluated and presented as graphs near the bedside. Thus, statistical routines can increase the better understanding of instable clinical conditions and lend support to the decision making process. The early detection of a pathological trend in a patient whose ability to compensate is still present provides necessary time for diagnostic or preventive countermeasures in case of emergency.

Computers

Histopathological criteria for progressive dementia disorders: clinical-pathological correlation and classification by multivariate data analysis.

Autopsied brains from 55 patients with dementia between 59-95 years of age (mean age 77.9 +/- 8.1 years) and 19 non-demented individuals between 46-91 years of age (mean age 74.3 +/- 10.5 years) were examined to establish histopathological criteria for normal ageing, primary degenerative [Alzheimer's disease (AD)/senile dementia of Alzheimer type (SDAT)] and vascular (multi-infarct) dementia (MID) disorders. Senile/neuritic plaques, neurofibrillary tangles, microscopic infarcts and perivascular serum protein deposits were quantified in the frontal lobe (Brodmann area 10) and in the hippocampus. The demented patients were classified according to the DSM-III criteria into AD/SDAT and MID. Operationally defined histopathological criteria for dementias, based on the degree/amount of the histopathological changes seen in aged non-demented patients, were postulated. The demented patients were clearly separable into three histopathological types, namely AD/SDAT, MID and AD-MID, the dementia type where both the degenerative and the vascular changes are coexistent in greater extent than are seen in the non-demented individuals. Using general clinical, gross neuroanatomical and histopathological data three separate dementia classes, namely AD/SDAT, MID and AD-MID, were visualized in two-dimensional space by multivariate data analysis. This analysis revealed that the pathology in the AD-MID patients was not merely a linear combination of the pathology in AD/SDAT and MID, indicating that AD-MID might represent a dementia type of its own. The clinical diagnosis for AD/SDAT and MID was certain in only half of the AD/SDAT and one third of the MID cases when evaluated histopathologically and by multivariate data analysis. AD/SDAT, MID and AD-MID were histopathologically diagnosed in 49%, 24% and 27%, respectively, of all the dementia cases studied. Opposite correlation between the number of tangles, plaques and the patient age in non-demented and AD/SDAT cases were observed, indicating that the pathogenesis of tangles and plaques in the two groups of patients might be different and that AD/SDAT might not be a form of an exaggerated ageing process.

Aged

[The effect of smoking habit on aortic pulse wave velocity using a new method for data analysis].

We measured aortic pulse wave velocity (PWV) in 168 male adult cases of various arteriosclerotic diseases. In order to evaluate the effects of age, smoking habits, alcohol intake, and blood pressure, we applied the least median of squares (LMS) regression which was considered to be very useful for data analysis. The results showed that PWV level increased with age. Furthermore smoking was associated with increasing PWV level and this effect was also related to age. We concluded that the PWV was valuable as an index of arteriosclerosis, and instead of the classical least squares method, LMS regression was very useful for analysis of medical data.

Adult

Diagnostic accuracy of pancreatic enzymes evaluated by use of multivariate data analysis.

We analyzed pancreatic enzyme data from 508 patients with suspected pancreatitis by neural network analysis, by an Expert multirule generation protocol, and by receiver-operator characteristic (ROC) curve analysis of a single test result. Neural network analysis showed that use of lipase provided the best means for diagnosing pancreatitis. Diagnostic accuracies achieved by using amylase only, lipase only, and amylase and lipase in combination were 76%, 82%, and 84%, respectively. Use of the Expert rule generation protocol provided a diagnostic accuracy of 92% when rules for single and multiple samplings were combined. ROC curve analysis for initial enzyme activities showed the maximal diagnostic accuracy to be 82% and 85% for amylase and lipase, respectively; use of peak enzyme activities yielded accuracies of 81% and 88%, respectively. The evaluation of laboratory test data should include analysis of the diagnostic accuracy of laboratory tests by multivariate techniques such as neural network analysis or an Expert systems approach. Multivariate analysis should allow for a more realistic assessment of the diagnosis accuracy of laboratory tests because all the available data are included in the evaluation.

Amylases

Implementing a training resource for large-scale genomic data analysis in the All of Us Researcher Workbench.

A lack of representation in genomic research and limited access to computational training create barriers for many researchers seeking to analyze large-scale genetic datasets. The All of Us Research Program provides an unprecedented opportunity to address these gaps by offering genomic data from a broad range of participants, but its impact depends on equipping researchers with the necessary skills to use it effectively. The All of Us Biomedical Researcher (BR) Scholars Program at Baylor College of Medicine aims to break down these barriers by providing early-career researchers with hands-on training in computational genomics through the All of Us Evenings with Genetics Research Program. The year-long program begins with the faculty summit, an in-person computational boot camp that introduces scholars to foundational skills for using the All of Us dataset via a cloud-based research environment. The genomics tutorials focus on genome-wide association studies (GWASs), utilizing Jupyter Notebooks and the Hail computing framework to provide an accessible and scalable approach to large-scale data analysis. Scholars engage in hands-on exercises covering data preparation, quality control, association testing, and result interpretation. By the end of the summit, participants will have successfully conducted a GWAS, visualized key findings, and gained confidence in computational resource management. This initiative expands access to genomic research by equipping early-career researchers from a variety of backgrounds with the tools and knowledge to analyze All of Us data. By lowering barriers to entry and promoting the study of representative populations, the program fosters innovation in precision medicine and advances equity in genomic research.

Humans

Proficiency of the Tradescantia-micronucleus image analysis system for scoring micronucleus frequencies and data analysis.

The Tradescantia-micronucleus (Trad-MCN) bioassay is an efficient short-term test for genotoxicity of pollutants. In order to increase the efficiency and to standardize the micronucleus (MCN) scoring process, an automated scoring system was developed using the principle of image analysis in computer science. This assemblage is called the Tradescantia-micronucleus image analysis (Trad-MCNIA) system. The MCN frequencies scored by this system were compared with those scored by human observation for its proficiency. A set of low MCN frequency (around 5 MCN/100 tetrads) slides prepared from a control group, a set of medium MCN frequency (around 20 MCN/100 tetrads) slides prepared from sodium azide treated plant cuttings and a set of high MCN frequency (around 50 MCN/100 tetrads) slides prepared from X-ray treated materials were used for this study. In the low MCN frequency slides, the Trad-MCNIA system scored about the same value as human observation. In the medium and high frequency slides, MCN frequencies scored by the system were lower than those scored by human observers. This discrepancy was corrected by increasing the power of the objective of the microscope in the system. The MCN frequencies scored by the system attained 90% congruity with those scored by human observers after the correction. The scoring speed of the system was about 3.5 times as fast as that by human observers, and the data could be statistically analyzed immediately after the data scores were recorded. Further improvements can be made by upgrading the video camera and the computer speed.

Azides

Recruitment in NHLBI population-based studies and randomized clinical trials: data analysis and survey results.

Data on screening and recruitment from current and previous NHLBI population-based studies (PBSs) and randomized clinical trials (RCTs) were examined. In only two of the studies examined was the projected recruitment completed within the planned recruitment period. The shape of the graph of the relation between enrollment of participants and time varies by study. A single summary statistic for measuring the efficiency of recruitment in RCTs and PBSs is proposed and applied to the examined studies. In addition to providing summary data on recruitment for several studies, this article reports the survey results of a questionnaire sent to the coordinating centers of currently and previously funded National Heart, Lung, and Blood Institute and Veteran's Administration studies. The purpose was to ascertain the desirability of recommending that a generic core of information be collected on recruitment and screening in future studies. Most respondents believed that comparing data collected uniformly and prospectively might be helpful in designing further studies. The variables most respondents believed to be potentially useful are described.

Clinical Trials as Topic

A procedure for data analysis of the rodent micronucleus test involving a historical control.

No standard procedure of data analysis for rodent micronucleus tests involving historical controls has been established. In the present paper, under the presumption that the distribution of the historical control is stable and reliable, a procedure with three statistical steps is proposed to analyze the frequency of micronucleated polychromatic erythrocytes (MNPCEs). In the first step, the frequencies of MNPCEs in negative and positive control groups of a current experiment of the micronucleus test are compared with the distribution of historical negative and positive controls to examine the technical validity of the current experiment. In the second step, the frequency of MNPCEs in each treatment group is compared with the distribution of the historical negative control. In the third step, the dose-response relation is tested with the Cochran-Armitage trend test. A Monte Carlo stimulation study shows that the power of this procedure is acceptable and also this procedure is robust. An application of this procedure on real data reveals that it is effective in detecting clastogenic chemicals when the probability of a type I error is nearly .01.

Animals

Secondary data analysis: research method for the clinical nurse specialist.

This article presents a description of secondary data analysis and suggests that this type of research methodology may be helpful in facilitating research by the clinical nurse specialist (CNS). The article discusses the advantages and disadvantages of the use of this method specifically in relation to the CNS and offers suggestions for sources of data.

Data Collection

Comparison of Ehrlich ascites tumour and mouse liver cells by analytical subcellular fractionation combined with a sensitive computational method for data analysis.

A simple method of analytical subcellular fractionation, combined with a sensitive computational method for data analysis and presentation, has been used to reinvestigate the distribution and relative amounts of several enzymes in the cytoplasmic and plasma membranes of two different cell types: one is a neoplastic, transformed cell type (Ehrlich ascites tumour cells), the other an untransformed, highly differentiated cell type (liver hepatocytes plus Kupffer and endothelial cells). In general the distribution of the enzymes in particular membranes is similar in the two cell types, however the relative amounts differ. Ehrlich ascites tumour cells have a higher specific activity of galactosyltransferase and ouabain-sensitive (Na,K)ATPase, while liver cells have higher glucose-6-phosphatase, 5'-nucleotidase and succinate dehydrogenase activity. These differences appear to be correlated with morphological and, in some cases, functional differences between the two cell types.

5'-Nucleotidase

A simplified method of echocardiographic data analysis.

Rapid accurate analysis of echocardiographic data is accomplished using a sonic digitizer and programmable calculator. This method allows the echocardiographer to select technically optimal areas of the recording for analysis. The resolution of the measuring device is 0.1 mm. A hardcopy printout of both measurement and calculation is provided. Instead of expensive on-line computer, an inexpensive programmable calculator is used.

Computers

A consultation system constructor for medical data analysis.

MAD is a system that helps an expert data analyst in a specific application domain (like epidemiology or image analysis) to build reasoning models aimed at fulfilling specific tasks. These models may be subsequently used to guide doctors in the analysis of a set of data referring to a specific ground domain. Expert knowledge is represented at various levels: a general description of an application domain and various models that formalize the reasoning followed to perform specific tasks within a defined application domain. Reasoning models are represented as rules of propositional calculus, and a meta-knowledge permits to support knowledge acquisition. During the consultation, different external programs may be run when needed, without the doctor having to learn how to use them. MAD is written in Golden Common LISP and may be linked to any external software for data analysis, provided it runs under MS-DOS and does not require more than 192 Kb. Examples of application of the system to epidemiology and image analysis are given.

Computer Simulation