PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Quantitative modeling and data analysis of SELEX experiments.

SELEX (systematic evolution of ligands by exponential enrichment) is an experimental procedure that allows the extraction, from an initially random pool of DNA, of those oligomers with high affinity for a given DNA-binding protein. We address what is a suitable experimental and computational procedure to infer parameters of transcription factor-DNA interaction from SELEX experiments. To answer this, we use a biophysical model of transcription factor-DNA interactions to quantitatively model SELEX. We show that a standard procedure is unsuitable for obtaining accurate interaction parameters. However, we theoretically show that a modified experiment in which chemical potential is fixed through different rounds of the experiment allows robust generation of an appropriate dataset. Based on our quantitative model, we propose a novel bioinformatic method of data analysis for such a modified experiment and apply it to extract the interaction parameters for a mammalian transcription factor CTF/NFI. From a practical point of view, our method results in a significantly improved false positive/false negative trade-off, as compared to both the standard information theory based method and a widely used empirically formulated procedure.

Animals↗

A statistical approach for array CGH data analysis.

BACKGROUND: Microarray-CGH experiments are used to detect and map chromosomal imbalances, by hybridizing targets of genomic DNA from a test and a reference sample to sequences immobilized on a slide. These probes are genomic DNA sequences (BACs) that are mapped on the genome. The signal has a spatial coherence that can be handled by specific statistical tools. Segmentation methods seem to be a natural framework for this purpose. A CGH profile can be viewed as a succession of segments that represent homogeneous regions in the genome whose BACs share the same relative copy number on average. We model a CGH profile by a random Gaussian process whose distribution parameters are affected by abrupt changes at unknown coordinates. Two major problems arise: to determine which parameters are affected by the abrupt changes (the mean and the variance, or the mean only), and the selection of the number of segments in the profile. RESULTS: We demonstrate that existing methods for estimating the number of segments are not well adapted in the case of array CGH data, and we propose an adaptive criterion that detects previously mapped chromosomal aberrations. The performances of this method are discussed based on simulations and publicly available data sets. Then we discuss the choice of modeling for array CGH data and show that the model with a homogeneous variance is adapted to this context. CONCLUSIONS: Array CGH data analysis is an emerging field that needs appropriate statistical tools. Process segmentation and model selection provide a theoretical framework that allows precise biological interpretations. Adaptive methods for model selection give promising results concerning the estimation of the number of altered regions on the genome.

Algorithms↗

Impact of rosuvastatin use on costs and outcomes in patients at high risk for cardiovascular disease in US managed care and medicare populations: A data analysis.

BACKGROUND: High blood cholesterol is a major modifiable risk factor for coronary heart disease (CHD) and stroke. OBJECTIVE: The aim of this study was to estimate the economic impact of rosuvastatin calcium use in patients at high risk for CHD and stroke, according to the National Cholesterol Education Program Adult Treatment Panel (ATP) III guidelines. METHODS: An economic simulation model was developed that used a Markov process to project the number of cardiovascular events and associated costs in a high-risk population in various treatment scenarios. According to the ATP III, high-risk patients are those with CHD, atherosclerosis of peripheral and/or cerebral arteries, diabetes, and/or multiple other risk factors conferring a risk of at least 20% within 10 years. Data on population characteristics and costs of cardiovascular disease (CVD) were obtained from claims data sets from employer-funded commercial and Medicare health plans in the United States. Treatment of lipid disorders was translated into CVD risk reduction based on results from the Heart Protection Study. The estimated efficacies of individual lipid-lowering drugs were based on data published in package inserts. The model generated costs at the health plan level of lipid-lowering therapy in high-risk patients and the number and total costs of cardiovascular events. Estimates were compared for scenarios representing the mix of treatments used before and after the introduction of rosuvastatin. Estimates were generated separately for commercial and Medicare health plans. RESULTS: For every 1 million members of a commercial health plan, an estimated 44,457 met ATP III criteria for high-risk status. Use of rosuvastatin in place of other 3-hydroxy-3-methylglutaryl coenzyme A reductase inhibitors ("statins") by 11 % of these patients over a period of 5 years was estimated to result in 36 fewer cardiovascular events and a net savings of US 4.03 million dollars. A Medicare plan of 1 million members with an estimated 433,268 high-risk patients and 7% rosuvastatin use was estimated to avoid 727 events and save US 34.32 million dollars. CONCLUSIONS: The results of this data analysis suggest that increasing the use of rosuvastatin can result in cardiovascular event reduction and cost savings. Because the impact of lipid-modifying therapy on cardiovascular risk has not been thoroughly documented in controlled clinical studies, our model assumed that incremental lipid changes had effects in proportion to the magnitude of change.

Adult↗

Data analysis of kinetic modelling used in drug stability studies: isothermal versus nonisothermal assays.

PURPOSE: Kinetic modelling was applied to predict the stability of cholecystokinin fragment CCK-4 in aqueous solution, which was analyzed by isothermal and nonisothermal methods using a validated stability indicating HPLC method. METHODS: The isothermal studies were performed in the temperature range 40 to 80 degrees C at pH 12 and ionic strength 0.01 M as constants, whereas nonisothermal stability studies were performed using a linear increasing temperature program, heating rate 0.25 degrees C/h and a temperature interval 40-82 degrees C. The isothermal studies require two-step linear regression to estimate the parameters, resulting in a well-defined confidence interval. Nonisothermal kinetic studies require nonlinear or linear regression by previous transformation of data to estimate the parameters. In this case, the two most popular approaches, derivative and integral, were used and compared. RESULTS: Under isothermal conditions, an apparent first-order degradation process was observed at all temperatures. The linear Arrhenius plot suggested that the CCK-4 degradation mechanism was the same within the studied temperature range, with quite large uncertainties due to the small number of degrees of freedom based only on the scatter in the plot, and giving an estimated shelf life at 25 degrees C of 35.2 days. The derivative approach yields high variability in the Arrhenius parameters, since they are dependent on the number of polynomial terms chosen, so several statistical criteria were applied to select the best model. The integral approach allows activation parameters to be calculated directly from experimental data, and provides results in good agreement with those of the traditional method, but have the advantage that the uncertainty in the final result directly reflects the goodness of fit of the experimental data to the chosen kinetic model. The application of the bootstrap technique to estimating confidence limits for the Arrhenius parameters and shelf life is also illustrated, and shows there is no difference between the asymptotic and bootstrap confidence intervals. CONCLUSIONS: Nonisothermal studies give us fast and valuable information about drug stability, although their potential for predicting isothermal behaviour is conditioned by the data analysis method applied.

Drug Stability↗

Assay development and data analysis of receptor-ligand binding based on scintillation proximity assay.

In this paper, we described the optimization of a generic binding assay to measure ligand-receptor interactions for peroxisome proliferator-activated receptors (PPARs). The assay is based on scintillation proximity assay, in which a protein is coated on scintillant-incorporated beads, and a radiolabeled ligand stimulates the beads to emit a signal by binding to the immobilized protein. An intrinsic binding affinity of unlabeled ligands is determined by competitive displacement of the radioligand. The protein coating and ligand binding are achieved in one step by simply mixing ligands, protein and beads in sequence. No additional steps of pre-coating and washing of beads are required. Protein is captured on beads effectively by electrostatic interactions, thus no affinity labeling of protein is required. In data analysis, ligands are grouped into two classes based on their binding affinities. For tight binding ligands, an equation is derived to accurately determine the binding affinity. Otherwise a general equation applies. This quantitative and high throughput assay provides a tool to screen a large library of molecules in search of potent ligands.

PPAR gamma↗

Performance of plate-based cytokine flow cytometry with automated data analysis.

BACKGROUND: Cytokine flow cytometry (CFC) provides a multiparameter alternative to ELISPOT assays for rapid quantitation of antigen-specific T cells. To increase the throughput of CFC assays, we have optimized methods for stimulating, staining, and acquiring whole blood or PBMC samples in 96-well or 24-well plates. RESULTS: We have developed a protocol for whole blood stimulation and processing in deep-well 24- or 96-well plates, and fresh or cryopreserved peripheral blood mononuclear cell (PBMC) stimulation and processing in conventional 96-well round-bottom plates. Samples from both HIV-1-seronegative and HIV-1-seropositive donors were tested. We show that the percent response, staining intensity, and cell recovery are comparable to stimulation and processing in tubes using traditional methods. We also show the equivalence of automated gating templates to manual gating for CFC data analysis. CONCLUSION: When combined with flow cytometry analysis using an automated plate loader and an automated analysis algorithm, these plate-based methods provide a higher throughput platform for CFC, as well as reducing operator-induced variability. These factors will be important for processing the numbers of samples required in large clinical trials, and for epitope mapping of patient responses.

Algorithms↗

Plurality and resemblance in fMRI data analysis.

We apply nine analytic methods employed currently in imaging neuroscience to simulated and actual BOLD fMRI signals and compare their performances under each signal type. Starting with baseline time series generated by a resting subject during a null hypothesis study, we compare method performance with embedded focal activity in these series of three different types whose magnitudes and time courses are simple, convolved with spatially varying hemodynamic responses, and highly spatially interactive. We then apply these same nine methods to BOLD fMRI time series from contralateral primary motor cortex and ipsilateral cerebellum collected during a sequential finger opposition study. Paired comparisons of results across methods include a voxel-specific concordance correlation coefficient for reproducibility and a resemblance measure that accommodates spatial autocorrelation of differences in activity surfaces. Receiver-operating characteristic curves show considerable model differences in ranges less than 10% significance level (false positives) and greater than 80% power (true positives). Concordance and resemblance measures reveal significant differences between activity surfaces in both data sets. These measures can assist researchers by identifying groups of models producing similar and dissimilar results, and thereby help to validate, consolidate, and simplify reports of statistical findings. A pluralistic strategy for fMRI data analysis can uncover invariant and highly interactive relationships between local activity foci and serve as a basis for further discovery of organizational principles of the brain. Results also suggest that a pluralistic empirical strategy coupled formally with substantive prior knowledge can help to uncover new brain-behavior relationships that may remain hidden if only a single method is employed.

Brain↗

Physiological consequences of experimental cerebral missile injury and use of data analysis to predict survival.

The authors describe cerebrovascular and cerebral metabolic changes in monkeys, subjected to cerebral missile injury. After injury with BB pellet at 90 m/sec, there is a rapid rise in intracranial pressure (ICP), which reaches a peak 2 to 5 minutes posttrauma, and then falls to about 20 to 30 mm Hg. This, with a fall in mean blood pressure (MBP), results in a 50% reduction in cerebral perfusion pressure (CPP), Cerebral blood flow (CBF) is also reduced, although acutely there is no close relationship with (CPP). Cerebrovascular resistance falls initially and then at 30 minutes rises to very high values. Cerebral metabolic rates (CMR's) for oxygen fall after injury and remain low for the rest of the animal's life; CMR's for lactate rise immediately after injury and persists for 5 hours, then fall. After injury with a faster missile (180 m/sec), the ICP rises higher and faster, and the peak is shorter. The CCP is reduced in this injury to approximately 30 mm Hg, and only one animal survived more than 1 hour. With the conventional forms of data analysis, the length of survival after injury correlates well with MBP, ICP, and CBF, but separately they were completely unsatisfactory for prediction of an individuals prognosis. With the technique of multiple linear regression analysis, the survival of individual animals could be predicted with great accuracy. This is possible also when two postinjury parameters,CBF and MBP, are used.

Animals↗

The role of multiple regression and exploratory data analysis in the development of leukemia incidence risk models for comparison of radionuclide air stack emissions from nuclear and coal power industries.

Risks associated with power generation must be identified to make intelligent choices between alternate power technologies. Radionuclide air stack emissions for a single coal plant and a single nuclear plant are used to compute the single plant leukemia incidence risk and total industry leukemia incidence risk. Leukemia incidence is the response variable as a function of radionuclide bone dose for the six proposed dose response curves considered. During normal operation a coal plant has higher radionuclide emissions than a nuclear plant and the coal industry has a higher leukaemia incidence risk than the nuclear industry, unless a nuclear accident occurs. Variation of nuclear accident size allows quantification of the impact of accidents on the total industry leukemia incidence risk comparison. The leukemia incidence risk is quantified as the number of accidents of a given size for the nuclear industry leukemia incidence risk to equal the coal industry leukemia incidence risk. The general linear model is used to develop equations that relate the accident frequency required for equal industry risks to the magnitude of the nuclear emission. Exploratory data analysis revealed that the relationship between the natural log of accident number versus the natural log of accident size is linear.

Journal Article↗

[Microarrays: technologies overview and data analysis].

DNA microarrays are a powerful tool to investigate differential gene expression for thousands of genes simultaneously. In this review, recent advances in DNA microarray technologies and their applications are examined. Various DNA microarray platforms are described along with their methods for fabrication and their use. In addition some algorithms and tools for the analysis of microarray expression data, including clustering methods, partitioning and machine learning methods are discussed.

Gene Expression Profiling↗

Geometrical validation of intravascular ultrasound radiofrequency data analysis (Virtual Histology) acquired with a 30 MHz boston scientific corporation imaging catheter.

Recently, the plaque characterization field was explored with the use of the substrate (frequency domain analysis) rather than the envelope (amplitude or gray-scale imaging) of the intravascular ultrasound (IVUS) radiofrequency data. However, there is no data about the agreement of quantitative outcome between the two methods. The aim of this study was to assess the correlation and agreement between quantitative coronary ultrasound and the geometrical measurements provided by the spectral analysis of ultrasound radiofrequency data [IVUS-Virtual Histology (IVUS-VH), Volcano Therapeutics). Twenty-five patients were included in this study. The IVUS catheter used was a commercially available mechanical sector scanner (Ultracross 2.9 Fr 30 MHz catheter, Boston Scientific) covered with an outer sheath. IVUS-VH significantly underestimated lumen [relative difference (RD)=14.8+/-5.6; P<0.001], vessel (RD=14.1+/-4.8; P<0.001), and plaque (RD=11.5+/-10.8; P<0.001) cross-sectional areas (CSAs). Nevertheless, when adjusted for the ultrasound propagation delay caused by the sheath, relative differences of measurements were remarkably low (0.49%+/-6.3%, P=0.64 for lumen; 2.33%+/-4.6%, P=0.007 for vessel; and 4.2%+/-10.4%, P=0.005 for plaque CSA). These data suggest that the volumetric output of the IVUS-VH software underestimates measurements when acquired with a 30 MHz catheter. However, after applying a mathematical adjustment method for the ultrasound propagation delay caused by the outer sheath of the 30 MHz catheter, relative differences of direct measurements were negligible. These results suggest that ultrasound radiofrequency data analysis could provide, aside from precise compositional data, an accurate geometrical output.

Cardiac Catheterization↗

An improved data analysis method for interleukin 2 microassay.

Development of the interleukin 2(IL 2) microassay, coupled with the use of highly purified or recombinant factors has allowed a detailed examination of the mechanism of action of this important biological response modifier. However, probit analysis of the microassay data does not allow inherent error of the system to be approximated nor can units of activity be assessed for significance. A computer program was developed to analyze the validity of each regression line and to generate 95% confidence intervals around each line. This program employs analysis of variance, linear regression analysis and the parallel line assay to fix confidence intervals for each IL 2 unit value. The use of recombinant IL 2 as an immunomodulator in clinical settings warrants a more precise statistical method to evaluate normal fluctuations of this factor than currently in use. The development of such a method is presented here.

Biological Assay↗

Need for optimal body composition data analysis using air-displacement plethysmography in children and adolescents.

Air-displacement plethysmography (ADP) is now widely used for body composition measurement in pediatric populations. However, the manufacturer's software developed for adults leaves a potential bias for application in children and adolescents, and recent publications do not consistently use child-specific corrections. Therefore we analyzed child-specific ADP corrections with respect to quantity and etiology of bias compared with adult formulas. An optimal correction protocol is provided giving step-by-step instructions for calculations. In this study, 258 children and adolescents (143 girls and 115 boys ranging from 5 to 18 y) with a high prevalence of overweight or obesity (28.0% in girls and 22.6% in boys) were examined by ADP applying the manufacturer's software as well as published equations for child-specific corrections for surface area artifact (SAA), thoracic gas volume (TGV), and density of fat-free mass (FFM). Compared with child-specific equations for SAA, TGV, and density of FFM, the mean overestimation of the percentage of fat mass using the manufacturer's software was 10% in children and adolescents. Half of the bias derived from the use of Siri's equation not corrected for age-dependent differences in FFM density. An additional 3 and 2% of bias resulted from the application of adult equations for prediction of SAA and TGV, respectively. Different child-specific equations used to predict TGV did not differ in the percentage of fat mass. We conclude that there is a need for child-specific equations in ADP raw data analysis considering SAA, TGV, and density of FFM.

Adolescent↗

GLUT4 protein expression in obese and lean 12-month-old rats: insights from different types of data analysis.

GLUT4 protein expression in white adipose tissue (WAT) and skeletal muscle (SM) was investigated in 2-month-old, 12-month-old spontaneously obese or 12-month-old calorie-restricted lean Wistar rats, by considering different parameters of analysis, such as tissue and body weight, and total protein yield of the tissue. In WAT, an approximately 70% decrease was observed in plasma membrane and microsomal GLUT4 protein, expressed as microg protein or g tissue, in both 12-month-old obese and 12-month-old lean rats compared to 2-month-old rats. However, when plasma membrane and microsomal GLUT4 tissue contents were expressed as g body weight, they were the same. In SM, GLUT4 protein content, expressed as microg protein, was similar in 2-month-old and 12-month-old obese rats, whereas it was reduced in 12-month-old obese rats, when expressed as g tissue or g body weight, which may play an important role in insulin resistance. Weight loss did not change the SM GLUT4 content. These results show that altered insulin sensitivity is accompanied by modulation of GLUT4 protein expression. However, the true role of WAT and SM GLUT4 contents in whole-body or tissue insulin sensitivity should be determined considering not only GLUT4 protein expression, but also the strong morphostructural changes in these tissues, which require different types of data analysis.

Adipose Tissue↗

Increasing efficiency and precision of data analysis: multivariate vs. univariate statistical techniques.

In deciding upon an appropriate analysis strategy during the planning phase of a research study, it is important to specify all of the independent and dependent variables to be included. From there, a technique should be chosen that will yield all of the information desired in the most efficient, precise, and powerful way. Frequently, the method of choice in nursing research will be a multivariate technique, since so many studies involve numerous variables whose effects on, or relationships with, other variables are of interest, as well as involving additional variables that should be taken into account for control or generalizability purposes. Not discussed in this paper, but worth mention, is the fact that all of these techniques involve various underlying assumptions (e.g., normality, homogeneity of variance, and independence in ANOVA) that must be met in order to use the techniques appropriately. If the assumptions are not met, the researcher might want to consider the use of nonparametric techniques--and there are both univariate and multivariate techniques available to choose from (Hollander & Wolfe, 1973; Siegel, 1956). The general advantages of a multivariate technique rather than separate univariate techniques would apply in terms of nonparametric statistics as well as parametric statistics. Increasing the accuracy, power, and efficiency of data analysis strategies should be a major concern to researchers.(ABSTRACT TRUNCATED AT 250 WORDS)

Analysis of Variance↗

Quantitative trait associated microarray gene expression data analysis.

Selection on phenotypes may cause genetic change. To understand the relationship between phenotype and gene expression from an evolutionary viewpoint, it is important to study the concordance between gene expression and profiles of phenotypes. In this study, we use a novel method of clustering to identify genes whose expression profiles are related to a quantitative phenotype. Cluster analysis of gene expression data aims at classifying genes into several different groups based on the similarity of their expression profiles across multiple conditions. The hope is that genes that are classified into the same clusters may share underlying regulatory elements or may be a part of the same metabolic pathways. Current methods for examining the association between phenotype and gene expression are limited to linear association measured by the correlation between individual gene expression values and phenotype. Genes may be associated with the phenotype in a nonlinear fashion. In addition, groups of genes that share a particular pattern in their relationship to phenotype may be of evolutionary interest. In this study, we develop a method to group genes based on orthogonal polynomials under a multivariate Gaussian mixture model. The effect of each expressed gene on the phenotype is partitioned into a cluster mean and a random deviation from the mean. Genes can also be clustered based on a time series. Parameters are estimated using the expectation-maximization algorithm and implemented in SAS. The method is verified with simulated data and demonstrated with experimental data from 2 studies, one clusters with respect to severity of disease in Alzheimer's patients and another clusters data for a rat fracture healing study over time. We find significant evidence of nonlinear associations in both studies and successfully describe these patterns with our method. We give detailed instructions and provide a working program that allows others to directly implement this method in their own analyses.

Animals↗

Biclustering algorithms for biological data analysis: a survey.

A large number of clustering approaches have been proposed for the analysis of gene expression data obtained from microarray experiments. However, the results from the application of standard clustering methods to genes are limited. This limitation is imposed by the existence of a number of experimental conditions where the activity of genes is uncorrelated. A similar limitation exists when clustering of conditions is performed. For this reason, a number of algorithms that perform simultaneous clustering on the row and column dimensions of the data matrix has been proposed. The goal is to find submatrices, that is, subgroups of genes and subgroups of conditions, where the genes exhibit highly correlated activities for every condition. In this paper, we refer to this class of algorithms as biclustering. Biclustering is also referred in the literature as coclustering and direct clustering, among others names, and has also been used in fields such as information retrieval and data mining. In this comprehensive survey, we analyze a large number of existing approaches to biclustering, and classify them in accordance with the type of biclusters they can find, the patterns of biclusters that are discovered, the methods used to perform the search, the approaches used to evaluate the solution, and the target applications.

Algorithms↗

Distance from the ostium as an independent determinant of coronary plaque composition in vivo: an intravascular ultrasound study based radiofrequency data analysis in humans.

AIMS: Relative plaque composition, more than its morphology alone, is thought to play a pivotal role in determining propensity to vulnerability. Thus, we investigated in vivo whether the distance from coronary ostium to plaque location independently affects plaque composition in humans. This may help explaining the recently reported non-uniform distribution of culprit lesions along the vessel in acute coronary syndromes. METHODS AND RESULTS: In 51 consecutive patients (45 men), aged 38-76 years (mean age: 58+/-10), a non-culprit vessel was investigated through spectral analysis of IVUS radiofrequency data (IVUS-Virtual Histology). The study vessel was the left anterior descending artery in 23 (45%) patients; the circumflex artery in nine (18%), and right coronary artery in 19 (37%). The overall length of the region of interest, subsequently divided into 10 mm segments, was 41.5+/-13 mm long (range: 30.2-78.4). No significant change was observed in terms of relative plaque composition along the vessel with respect to fibrous, fibrolipidic, and calcified tissue, whereas the percentage of lipid core resulted to be increased in the first (median: 8.75%; IQR: 5.7-18) vs. the third (median: 6.1%; IQR: 3.2-12) (P=0.036) and fourth (median: 4.5%; IQR: 2.4-7.9) (P=0.006) segment. At multivariable regression analysis, distance from the ostium resulted to be an independent predictor of relative lipid content [beta=-0.28 (95%CI: -0.15, -0.41)], together with older age, unstable presentation, no use of statin, and presence of diabetes mellitus. CONCLUSION: Plaque distance from the coronary ostium, as an independent determinant of relative lipid content, is potentially associated to plaque vulnerability in humans.

Adult↗