PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Interpretation, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

An overview of relations among causal modelling methods.

This paper provides a brief overview to four major types of causal models for health-sciences research: Graphical models (causal diagrams), potential-outcome (counterfactual) models, sufficient-component cause models, and structural-equations models. The paper focuses on the logical connections among the different types of models and on the different strengths of each approach. Graphical models can illustrate qualitative population assumptions and sources of bias not easily seen with other approaches; sufficient-component cause models can illustrate specific hypotheses about mechanisms of action; and potential-outcome and structural-equations models provide a basis for quantitative analysis of effects. The different approaches provide complementary perspectives, and can be employed together to improve causal interpretations of conventional statistical results.

Causality↗

hprt mutant frequencies, nonpulmonary malignancies, and domestic radon exposure: "postmortem" analysis of an interesting hypothesis.

The hypothesis that exposure to domestic radon raises the risk for leukemia and other nonpulmonary cancers has been proposed and tested in a number of epidemiologic studies over the past decade. During this period, interest in this hypothesis was heightened by evidence of increased frequencies of mutations at the hypoxanthine guanine phosphoribosyl transferase (hprt) gene in persons exposed to domestic radon (Bridges BA et al. [1991]: Lancet 337:1187-1189). An extension of this study (Cole J et al. [lsqb[1996]: Radiat Res 145:61-69) and two independent studies (Albering HJ et al. [1992[: Lancet 340:739; Albering HJ et al. [1994[: Lancet 344:750-751) found that hprt mutant frequency was not correlated with domestic radon exposure, and two well-designed epidemiologic studies showed no evidence of a relation between radon exposure and leukemia in children or adults. In this report, we present additional data from a study of Colorado high school students showing no correlation between domestic radon exposure and hprt mutant frequency. We use reanalyses of previous studies of radon and hprt mutant frequency to identify problems with this assay as a biomarker for domestic radon exposure and to illustrate difficulties in interpreting the statistical data. We also show with analyses of combined data sets that there is no support for the hypothesis that domestic radon exposure elevates hprt mutant frequency. Taken together, the scientific evidence provides a useful example of the problems associated with analyzing and interpreting data that link environmental exposures, biomarkers, and diseases in epidemiologic studies.

Adolescent↗

Statistical mechanical approach for predicting the transition to non-B DNA structures in supercoiled DNA.

Supercoiling causes global twist of DNA structure and the supercoiled state has wide influence on conformational transition. A statistical mechanical approach was made for prediction of the transition probability to non-B DNA structures under torsional stress. A conditional partition function was defined as the sum over all possible states of the DNA sequence with basepair 1 and basepair n being in B-form helix and a recurrence formula was developed which expressed the partition function for basepair n with those for less number of pairs. This new definition permits a quick enumeration of every configuration of secondary structures. Energetic parameters of all conformations concerned, involving B-form, interior loop, cruciform and Z-form, were included in the equation. The probability of transition to each non-B conformation could be derived from these conditional partition functions. For treatment of effects of superhelicity, supercoiling energy was considered, and a twist of each conformation was determined to minimize the supercoiling energy. As the twist itself affects the transition probability, the whole scheme of equations was solved by renormalization technique. The present method permits a simultaneous treatment of several types of conformations under a common torsional stress. A set of energetic parameters of DNA secondary structures has been chosen for calculation. Some DNA sequences were submitted to the calculation, and all the sequences that we submitted gave stable convergence. Some of them have been investigated the critical supercoil density for the transition to non-B DNA structures. Even though the reliability of the set of parameters was not enough, the prediction of secondary structure transition showed good agreement with reported observation. Hence, the present algorithm can estimate the probability of local conformational change of DNA under a given supercoil density, and also be employed to predict some specific sequences in which conformational change is sensitive to superhelicity.

Algorithms↗

ROC and confusion analysis of structure comparison methods identify the main causes of divergence from manual protein classification.

BACKGROUND: Current classification of protein folds are based, ultimately, on visual inspection of similarities. Previous attempts to use computerized structure comparison methods show only partial agreement with curated databases, but have failed to provide detailed statistical and structural analysis of the causes of these divergences. RESULTS: We construct a map of similarities/dissimilarities among manually defined protein folds, using a score cutoff value determined by means of the Receiver Operating Characteristics curve. It identifies folds which appear to overlap or to be "confused" with each other by two distinct similarity measures. It also identifies folds which appear inhomogeneous in that they contain apparently dissimilar domains, as measured by both similarity measures. At a low (1%) false positive rate, 25 to 38% of domain pairs in the same SCOP folds do not appear similar. Our results suggest either that some of these folds are defined using criteria other than purely structural consideration or that the similarity measures used do not recognize some relevant aspects of structural similarity in certain cases. Specifically, variations of the "common core" of some folds are severe enough to defeat attempts to automatically detect structural similarity and/or to lead to false detection of similarity between domains in distinct folds. Structures in some folds vary greatly in size because they contain varying numbers of a repeating unit, while similarity scores are quite sensitive to size differences. Structures in different folds may contain similar substructures, which produce false positives. Finally, the common core within a structure may be too small relative to the entire structure, to be recognized as the basis of similarity to another. CONCLUSION: A detailed analysis of the entire available protein fold space by two automated similarity methods reveals the extent and the nature of the divergence between the automatically determined similarity/dissimilarity and the manual fold type classifications. Some of the observed divergences can probably be addressed with better structure comparison methods and better automatic, intelligent classification procedures. Others may be intrinsic to the problem, suggesting a continuous rather than discrete protein fold space.

Algorithms↗

MIDAW: a web tool for statistical analysis of microarray data.

MIDAW (microarray data analysis web tool) is a web interface integrating a series of statistical algorithms that can be used for processing and interpretation of microarray data. MIDAW consists of two main sections: data normalization and data analysis. In the normalization phase the simultaneous processing of several experiments with background correction, global and local mean and variance normalization are carried out. The data analysis section allows graphical display of expression data for descriptive purposes, estimation of missing values, reduction of data dimension, discriminant analysis and identification of marker genes. The statistical results are organized in dynamic web pages and tables, where the transcript/gene probes contained in a specific microarray platform can be linked (according to user choice) to external databases (GenBank, Entrez Gene, UniGene). Tutorial files help the user throughout the statistical analysis to ensure that the forms are filled out correctly. MIDAW has been developed using Perl and PHP and it uses R/Bioconductor languages and routines. MIDAW is GPL licensed and freely accessible at http://muscle.cribi.unipd.it/midaw/. Perl and PHP source codes are available from the authors upon request.

Algorithms↗

Optimal Control of Directional False Discovery Rates in Large-Scale Testing.

The high-throughput biomedical technology enables measurement of thousands of gene expression levels contemporaneously. A major task in analyzing these gene expression data is to identify both over-expressed and under-expressed genes. The popular two-group models select the non-null genes without further classifying them as overexpression or underexpression. Consequently, two-group decision rules are unable to constrain the numbers of falsely discovered over-expressed or under-expressed genes respectively. We propose a general three-group model that allows dependence between the test statistics and develop a decision rule that separately controls the two types of false discoveries. We show that the optimal decision rule in our three-group model has a special monotonic structure. By making use of this monotonic structure, we can linearize the two-directional false discovery rate constraints. We prove that our decision rule optimizes the expected number of true discoveries while controlling the proportions of falsely discovered over-expressed and under-expressed genes at desired levels simultaneously. The data-driven versions of the proposed procedures are suggested, and their consistency is established. Comparisons with state-of-the-art approaches and applications to genomic studies show that our procedures work well.

Humans↗

Systematics of New World monkeys (Platyrrhini, Primates) based on 16S mitochondrial DNA sequences: a comparative analysis of different weighting methods in cladistic analysis.

In order to investigate the effects of different weighting methods on a phylogeny reconstruction based on DNA sequences and to evaluate the phylogenetic information content of various secondary structures, a fragment of the large ribosomal mitochondrial gene (16S) was sequenced from 13 species of New World monkeys, three species of catarrhines, and Tarsius. The data were analyzed cladistically without weighting characters or changes, and with different weighting methods: a priori differential weights for transitions and transversions, two variants of dynamic weighting for each kind and direction of change, and successive approximations, using both the character consistency index (CI) and the rescaled consistency index (RC). The results were compared with published trees constructed from nuclear sequences of E-globins and morphological characters by different authors. The result of the analysis of the mtDNA data set with successive approximations, using the RC as weighting function, was the closest to the topology on which all molecular and morphological trees concur. Other relationships were unique to this tree. "Loops" were the type of secondary structure that showed maximum variation in sequence length and sites with the lowest character CI and RC. A large number of sites within loops showed high values for these indices, however, which suggests that uniform downweighting of these regions represents a large loss of phylogenetic information. Successive weighting, which assigns a weight for each particular character, seems to be a desirable alternative to this practice. We propose a new variant of dynamic weighting, which we call homoplasy-correcting dynamic weighting, that like dynamic weighting, is applicable to any kind of sequence, coding or noncoding.

Animals↗

Methodological considerations in rat brain BOLD contrast pharmacological MRI.

RATIONALE AND OBJECTIVES: Blood oxygen level dependent (BOLD) contrast pharmacological magnetic resonance imaging (phMRI) is an increasingly popular technique that allows the non-invasive investigation of spatial and temporal changes in rat brain function in response to pharmacological stimulation in vivo. Rat brain BOLD contrast phMRI is, at present, established in few neuropharmacological laboratories, and various issues associated with the technique require attention. The present review is primarily aimed at psychopharmacologists with no previous experience of phMRI, who are interested in the practical aspects that phMRI studies entail. RESULTS AND DISCUSSION: Experimental and analytical considerations, including anaesthesia, physiological monitoring, drug dose and delivery, scanning protocols, statistical approaches and the interpretation of phMRI data, are discussed.

Anesthesia↗

Inverse diffusion methods for data peak separation.

Previous methods for separation of overlapping data peaks include geometrical assessment and Fourier deconvolution. On the basis of inverse diffusion theory, we present new separation methods suitable for convenient programming and rapid calculation of Gaussian area contributions. Both continuum and discrete inverse diffusion models are described. Example computations are given for biological data: density gradient centrifugation, isoelectric focusing electrophoresis, and high-pressure liquid chromatography.

Centrifugation, Density Gradient↗

Molecular sequence accuracy: analysing imperfect data.

Molecular sequences are experimentally derived data that can be expected to contain errors as a result of diverse phenomena such as biological variation, molecular cloning artifacts, imperfect sequence determination, and data handling during contig assembly. Errors will affect the reliability of database searches and sequence alignments, but their impact may be minimized by the use of analytical techniques that anticipate that the data will be imperfect.

Amino Acid Sequence↗

Analysis of beat-to-beat cardiovascular hemodynamic variables obtained from long-term biotelemetry.

Manual methods of large volume data storage, retrieval, and analysis are difficult, time consuming, and present numerous opportunities for calculation errors. We have designed and implemented a comprehensive computer-based system for performing these functions. Development of this system was necessary since left ventricular (LV) blood pressure and two regional LV wall thickness measurements were obtained during long-term extracorporeal biotelemetry of miniswine for 24-h periods. During a single recording period over 100,000 individual cardiac cycles were recorded on analog tape and later analysed for determination of global myocardial oxygen demand and regional myocardial function. In addition, custom designed software was developed to determine the extent and duration of myocardial dysfunction. Batch file commands enabled the customized software to operate without prompting by the user thus optimizing the time usage of the computer, and the computer based data acquisition and analysis system. Although this system was designed specifically for analysing cardiovascular hemodynamic variables, it is flexible and can be applied to other experimental applications.

Analog-Digital Conversion↗