PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Audit in infection control--data analysis and infections in the intensive care unit.

The importance of high standards of infection control practices in intensive care units is widely accepted. Despite this, few objective measures of the efficacy of such practices have been developed. This hampers clinical audit, which is essential if nosocomial infections are to be kept to a minimum. Small sample size, multiple dependent variables, difficulty in matching populations, and the covert effects of confounding factors and bias, all add to the problems of designing meaningful audit programmes in intensive care settings. Nevertheless, useful analysis can be attempted now that appropriate standards and scoring systems are available. Computer systems can simplify data collection, analysis and reporting but they are no substitute for careful preparation and good statistical technique.

Computers↗

Retriever and CompareTable, two informatics tools for data analysis of high-density oligonucleotide arrays.

High-density oligonucleotide arrays are widely employed for detecting global changes in gene expression profiles of cells or tissues exposed to specific stimuli. Presented with large amounts of data, investigators can spend significant amounts of time analyzing and interpreting this array data. In our application of GeneChip arrays to analyze changes in gene expression in viral-infected epithelium, we have needed to develop additional computational tools that may be of utility to other investigators using this methodology. Here, I describe two executable programs to facilitate data extraction and multiple data point analysis. These programs run in a virtual DOS environment on Microsoft Windows 95/98/2K operating systems on a desktop PC. Both programs can be freely downloaded from the BioTechniques Software Library (www.BioTechniques.com). The first program, Retriever, extracts primary data from an array experiment contained in an Affymetrix textfile using user-inputted individual identification strings (e.g., the probe set identification numbers). With specific data retrieved for individual genes, hybridization profiles can be examined and data normalized. The second program, CompareTable, is used to facilitate comparison analysis of two experimental replicates. CompareTable compares two lists of genes, identifies common entries, extracts their data, and writes an output text file containing only those genes present in both of the experiments. The output files generated by these two programs can be opened and manipulated by any software application recognizing tab-delimited text files (e.g., Microsoft NotePad or Excel).

Computational Biology↗

OmicsQ: a user-friendly platform for interactive quantitative omics data analysis.

MOTIVATION: High-throughput omics technologies generate complex datasets with thousands of features that are quantified across multiple experimental conditions, but often suffer from incomplete measurements, missing values, and individually fluctuating variances. This requires analytical tools for accurate, deep and insightful biological interpretation, capable of dealing with a large variety of data properties and different amounts of completeness. Software capable of handling such data complexity and integrating with external applications for downstream analysis remains rare and mostly relies on programming-based environments, limiting accessibility for researchers without computational expertise. RESULTS: We present OmicsQ, an interactive, web-based platform designed to streamline quantitative omics data analysis. OmicsQ provides an intuitive, browser-based visualization interface that integrates established statistical processing tools. Those include robust batch correction, automated experimental design annotation, and handling of missing data without imputation, which maintains data integrity and avoids artifacts from a priori assumptions. OmicsQ seamlessly interacts with external applications (e.g. PolySTest, VSClust, ComplexBrowser) for statistical testing, clustering, analysis of protein complex behavior, and pathway enrichment, offering a comprehensive and flexible workflow from data import to biological interpretation that is broadly applicable across domains. AVAILABILITY AND IMPLEMENTATION: OmicsQ is implemented in R and Shiny and is available at https://computproteomics.bmb.sdu.dk/app_direct/OmicsQ. Source code and installation instructions: https://github.com/computproteomics/OmicsQ, DOI: 10.5281/zenodo.17778420.

Software↗

Q-RT-PCR: data analysis software for measurement of gene expression by competitive RT-PCR.

MOTIVATION: We have developed software to assist in the computation of gene expression from data obtained in competitive reverse transcription-polymerase chain reaction (RT-PCR). This report describes the mathematical basis of competitive RT-PCR and discusses the criteria which must be met to permit accurate estimations of gene expression to be obtained using this technique. RESULTS: The software that has been developed assists in both the assessment of assay performance (specifically in establishing the equality of amplification efficiency of the native and competitor templates) and in the routine analysis of data obtained in quantitation of gene expression by competitive RT-PCR. The software is a 100 kb module which functions as a Microsoft Excel add-in. It is compatible with both Windows and Mac versions of Excel 5 and Excel 7 on the Windows 95 platform, and employs the spreadsheet, statistical and graphing capabilities incorporated into Excel. AVAILABILITY: The software can be downloaded from http://www.grad.ttuhsc.edu/archive/. A brief summary in both HTML and Microsoft Word 6 format of the installation and use of the software is also located at this website.

Algorithms↗

Analysis of human serum by liquid chromatography-mass spectrometry: improved sample preparation and data analysis.

Discovery of biomarkers is a fast developing field in proteomics research. Liquid chromatography coupled on line to mass spectrometry (LC-MS) has become a powerful method for the sensitive detection, quantification and identification of proteins and peptides in biological fluids like serum. However, the presence of highly abundant proteins often masks those of lower abundance and thus generally prevents their detection and identification in proteomics studies. To perform future comparative analyses of samples from a serum bank of cervical cancer patients in a longitudinal and cross-sectional manner, methodology based on the depletion of high-abundance proteins followed by tryptic digestion and LC-MS has been developed. Two sample preparation methods were tested in terms of their efficiency to deplete high-abundance serum proteins and how they affect the repeatability of the LC-MS data sets. The first method comprised depletion of human serum albumin (HSA) on a dye ligand chromatographic and immunoglobulin G (IgG) on an immobilized Protein A support followed by tryptic digestion, fractionation by cation-exchange chromatography, trapping on a C18 column and reversed-phase LC-MS. The second method included depletion of the six most abundant serum proteins based on multiple immunoaffinity chromatography followed by tryptic digestion, trapping on a C18 column and reversed-phase LC-MS. Repeatability of the overall procedures was evaluated in terms of retention time and peak area for a selected number of endogenous peptides showing that the second method, besides being less time consuming, gave more repeatable results (retention time: <0.1% RSD; peak area: <30% RSD). Application of an LC-MS component detection algorithm followed by principal component analysis (PCA) enabled discrimination of serum samples that were spiked with horse heart cytochrome C from non-spiked serum and the detection of a concentration trend, which correlated to the amount of spiked horse heart cytochrome C to a level of 5 pmol cytochrome C in 2 microl original serum.

Animals↗

A new clustering method for microarray data analysis.

A novel clustering approach is introduced to overcome data missing and inconsistency of gene expression levels under different conditions in the stage of clustering. It is based on the so-called smooth score, which is defined for measuring the deviation of the expression level of a gene and the average expression level of all the genes involved under a condition. We present an efficient greedy algorithm for finding clusters with smooth score below a threshold after studying its computational complexity. The algorithm was tested intensively on random matrixes and a yeast data. It was shown to perform well in finding co-regulation patterns in a test with the yeast data.

Algorithms↗

Sparse data analysis.

In recent years there has been a growing interest in techniques capable of analyzing sparse data, particularly gathered during Phase III clinical trials, and there is now pressure on manufacturers to obtain more kinetic and dynamic information from Phase III studies. Techniques for the analysis of sparse data are reviewed drawing on a number of examples taken from pharmacokinetic and pharmacodynamic experiments.

Animals↗

Estimation of neocortical serotonin-2 receptor binding potential by single-dose fluorine-18-setoperone kinetic PET data analysis.

UNLABELLED: Because it satisfies most of the characteristics required to quantify in vivo neocortical serotonin-2 (5HT2) receptors, 18F-setoperone was selected for use in PET estimation of the neocortical 5HT2 binding parameters in baboons according to a single-dose paradigm. METHODS: The neocortical binding potential (i.e., Bmax/KD or the k3/k4 ratio) was assessed by three different methods, with the cerebellum taken as the reference structure in all instances. Method 1 was based on a Logan-Patlak graphical analysis of both cerebellar and neocortical data, which allows estimation of the neocortical k3'/k4 ratio; it required a separate estimation of k5 and k6 from classical nonlinear least-squares (NLSQ) three-compartment modeling of cerebellar data. Method 2 was an original combination of a four-compartment Logan-Patlak procedure for neocortical data and an NLSQ three-compartment procedure for cerebellar data, allowing the neocortical k3/k4 ratio to be obtained directly. In Method 3, an NLSQ three-compartment procedure was applied to cerebellar data and an NLSQ four-compartment procedure to neocortical data, allowing separate determinations of k3 and k4 for the neocortex and, in turn, the k3/k4 ratio. RESULTS: In all three methods, the arterial plasma input function was corrected for the presence of 18F-metabolites, and the vascular fraction was either fitted or fixed. Statistical analysis showed no significant difference among the k3/k4 values obtained from the three methods. Method 3 was the least stable because of an occasional poor NLSQ four-compartment fit on neocortical data. Method 2 provided the least cumbersome estimate of the k3/k4 ratio and was found easy and accurate for generating parametric maps of the 5HT2 binding potential. CONCLUSION: This method might be useful in clinical investigations to provide quantitative assessment of receptor binding potential. In semiquantitative investigations, the neocortical-to-cerebellum pseudoequilibrium ratio may be adequate, as suggested by the significant correlations with measured k3/k4 ratios found here.

Animals↗

Dependent masking and system life data analysis: Bayesian inference for two-component systems.

Data from field operations of a system is often used to estimate the reliability of components. Under ideal circumstances, this system field data contains the time to failure along with information on the exact component responsible for the system failure. However, in many cases, the exact component causing the failure of the system cannot be identified, and is considered to be masked. Previously developed models for estimation of component reliability from masked system life data have been based upon the assumption that masking occurs independently of the true cause of system failure. In this paper we develop a Bayesian methodology for estimating component reliabilities from masked system life data when the probability of masking is dependent upon the true cause of system failure. The Bayesian approach is illustrated for the case of a two-component system of exponentially distributed components.

Bayes Theorem↗

[Documentation of the surgical report with graphic statistical data analysis--a simplification of daily routine work].

A data collection system on microcomputer connected with an automatic medical report system for operations, was developed to facilitate both medical report as well as documentation. Linking different commercial software products by use of a Pascal programme, we were able to speed up daily routine work as well as establish efficient graphical statistics of patient data.

Cesarean Section↗

Findings in surgery for chronic otitis media. A retrospective data-analysis of 2225 cases followed for 2 years.

Data on middle ear disease and follow-up was recorded and analysed in 2225 ears operated on for chronic otitis media between 1958 and 1975. The closure rate of tympanic membrane perforations (70-80%) was not influenced by the presence of otorrhoea, but it was significantly lower in patients aged below 15 or over 40 yr and especially when tympanic membrane homograft had been used. For all types of tympanoplasty the mean improvement in air-conduction threshold was most marked at the lower frequencies, being some 10 dB in Type I, 15 dB in Type II and 8 dB in Type III and IV and it decreased by some 3 dB per octave. The combination of cholesteatoma and tympanosclerosis appeared to be rare. The post-operative incidence of otorrhoea was 10-15% after modified radical mastoidectomy and 20-25% after radical mastoidectomy.

Chronic Disease↗

The role of proxy information in missing data analysis.

This article investigates the role of proxy data in dealing with the common problem of missing data in clinical trials using repeated measures designs. In an effort to avoid the missing data situation, some proxy information can be gathered. The question is how to treat proxy information, that is, is it always better to utilize proxy information when there are missing data? A model for repeated measures data with missing values is considered and a strategy for utilizing proxy information is developed. Then, simulations are used to compare the power of a test using proxy to simply utilizing all available data. It is concluded that using proxy information can be a useful alternative when such information is available. The implications for various clinical designs are also considered and a data collection strategy for efficiently estimating parameters is suggested.

Canada↗

Interferometric data analysis based on Markov nonlinear filtering methodology.

For data processing in conventional phase shifting interferometry, Fourier transform, and least-squares-fitting techniques, a whole interferometric data series is required. We propose a new interferometric data processing methodology based on a recurrent nonlinear procedure. The signal value is predicted from the previous step to the next step, and the prediction error is used for nonlinear correction of an a priori estimate of the parameters phase, visibility, or frequency of interference fringes. Such a recurrent procedure is correct on the condition that the noise component be a Markov stochastic process realization. The accuracy and stability of the recurrent Markov nonlinear filtering algorithm were verified by computer simulations. It was discovered that the main advantages of the proposed methodology are dynamic data processing, phase error minimization, and high noise immunity against the influence of non-Gaussian noise correlated with the signal and the automatic solution of the phase unwrapping problem.

Algorithms↗

Multimodal CustOmics: A unified and interpretable multi-task deep learning framework for multimodal integrative data analysis in oncology.

Characterizing cancer presents a delicate challenge as it involves deciphering complex biological interactions within the tumor's microenvironment. Clinical trials often provide histology images and molecular profiling of tumors, which can help understand these interactions. Despite recent advances in representing multimodal data for weakly supervised tasks in the medical domain, achieving a coherent and interpretable fusion of whole slide images and multi-omics data is still a challenge. Each modality operates at distinct biological levels, introducing substantial correlations between and within data sources. In response to these challenges, we propose a novel deep-learning-based approach designed to represent multi-omics & histopathology data for precision medicine in a readily interpretable manner. While our approach demonstrates superior performance compared to state-of-the-art methods across multiple test cases, it also deals with incomplete and missing data in a robust manner. It extracts various scores characterizing the activity of each modality and their interactions at the pathway and gene levels. The strength of our method lies in its capacity to unravel pathway activation through multimodal relationships and to extend enrichment analysis to spatial data for supervised tasks. We showcase its predictive capacity and interpretation scores by extensively exploring multiple TCGA datasets and validation cohorts. The method opens new perspectives in understanding the complex relationships between multimodal pathological genomic data in different cancer types and is publicly available on Github.

Deep Learning↗

Computer simulation and data analysis of effector-target interactions: the extraction of binding parameters from effector and target conjugate frequencies data by using linear and nonlinear data-fitting transformations.

Binding isotherms for effector-target conjugation when effector conjugate frequencies are measured by holding constant the number of effector cells and by varying the number of target cells are characterized by two parameters, the maximum effector conjugate frequency, alpha max, and gamma, which is related to the dissociation constant of the conjugates formed, K d. The suitability of four linear transformations of these binding isotherms, as well as nonlinear data-fitting techniques, to provide estimates of alpha max and gamma is discussed. The strength and weakness of these procedures were investigated by calculating alpha max and gamma from different sets of 100 or 500 replicate "experiments," which were generated by using an algorithm that provides noise contributions to the conjugate frequencies with gaussian distributed errors. Both unweighted and weighted data points were used in these calculations. A similar analysis can also be performed for binding isotherms in which target conjugate frequencies are measured at different values of effector cells by holding constant the number of target cells. In this case, the binding isotherms are characterized by two parameters, the maximum target conjugate frequency, beta max, and delta, which is also related to K d. The results obtained demonstrate that if the experimental conditions are chosen properly, linear transformations and nonlinear fitting techniques provide reliable estimates for the binding parameters. Not all procedures, however, provide estimates with the same accuracy, and special emphasis to this fact must be given if the binding assays are performed at low values of the number of effector cells.

Algorithms↗

Determination of wheat quality by mass spectrometry and multivariate data analysis.

Multivariate analysis has been applied as support to proteome analysis in order to implement an easier and faster way of data handling based on separation by matrix-assisted laser desorption/ionisation time-of-flight mass spectrometry. The characterisation phase in proteome analysis by means of simple visual inspection is a demanding process and also insecure because subjectivity is the controlling element. Multivariate analysis offers, to a considerable extent, objectivity and must therefore be regarded as a neutral way to evaluate results obtained by proteome analysis. Proteome analysis of storage proteins from the wheat gluten complex based on two-dimensional electrophoresis and analysis of the N-terminal sequence has revealed a protein homologous to gamma-gliadins, tentatively associated with quality and within the molecular weight range 27-35 kDa. Further examinations of gliadin data based on mass spectrometry revealed that quality among wheat varieties could be determined by means of principal component analysis. Further examinations by interval partial least squares made it possible to encircle an overall optimal molecular weight interval from 31.5 to 33.7 kDa. The use of multivariate analysis on data from mass spectrometry has thus shown to be a promising technique to minimize the number of two-dimensional gels within the field of proteome analysis.

Gliadin↗