Statistical process control charts.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The mathematical and statistical evaluation of environmental data gains an increasing importance in environmental chemistry as the data sets become more complex. It is inarguable that different mathematical and statistical methods should be applied in order to compare results and to enhance the possible interpretation of the data. Very often several aspects have to be considered simultaneously, for example, several chemicals entailing a data matrix with objects (rows) and variables (columns). In this paper a data set is given concerning the pollution of 58 regions in the state of Baden-Württemberg, Germany, which are polluted with metals lead, cadmium, zinc, and with sulfur. For pragmatic reasons the evaluation is performed with the dichotomized data matrix. First this dichotomized 58 x 13 data matrix is evaluated by the Hasse diagram technique, a multicriteria evaluation method which has its scientific origin in Discrete Mathematics. Then the Partially Ordered Scalogram Analysis with Coordinates (POSAC) method is applied. It reduces the data matrix in plotting it in a two-dimensional space. A small given percentage of information is lost in this method. Important priority objects, like maximal and minimal objects (high and low polluted regions), can easily be detected by Hasse diagram technique and POSAC. Two variables attained exceptional importance by the data analysis shown here: TLS, Sulfur found in Tree Layer, is difficult to interpret and needs further investigations, whereas LRPB, Lead in Lumbricus Rubellus, seems to be a satisfying result because the earthworm is commonly discussed in the ecotoxicological literature as a specific and highly sensitive bioindicator.
Identifying and characterizing the structure in genome sequences is one of the principal challenges in modern molecular biology, and comparative genomics offers a powerful tool. In this paper, we introduce a hidden Markov model that allows a comparative analysis of multiple sequences related by a phylogenetic tree, and we present an efficient method for estimating the parameters of the model. The model integrates structure prediction methods for one sequence, statistical multiple alignment methods, and phylogenetic information. This unified model is particularly useful for a detailed characterization of DNA sequences with a common gene. We illustrate the model on a variety of homologous sequences.
Despite the wide availability of statistical programs designed to deal with longitudinal data from a multilevel perspective, many applied researchers remain unfamiliar with the benefits of this methodology, particularly for the evaluation of interventions. The authors present an example of multilevel modeling as part of the analysis of evaluation data from an HIV intervention study. Strategies for understanding multilevel models using longitudinal (panel) data are demonstrated and discussed. The authors illustrate how multiple linear regression models provide a convenient conceptual background to understanding how hierarchical linear models can be developed and interpreted. Multilevel analysis results are compared and contrasted with typical approaches through general linear models for repeated-measures data. Analyses are presented using the SPSS and HLM 5 software.
Computerized interactive 3-dimensional graphical displays were originally developed to aid in the exploration of multidimensional data from particle physics experiments. This technique can be equally well applied to speed the analysis of multivariate pharmaco-kinetic and pharmacodynamic data. The application of this technique to the results of a drug study demonstrated its effectiveness in providing a rapid overview of the data and in displaying new perspectives on the multivariate data, helping to identify sources of variability within the study. Multiple concentration-time curves from a given administration period can be distinguished within a single plot, using visual cues provided by rotating the display. Simultaneous comparison of large numbers of curves allows rapid evaluation of intersubject variability. Comparing the concentration-time curves from a single subject, each in succession, quickly identifies sources of intra-subject variability. Manipulating the display by a fourth dimensional parameter shows the degree of relationship between the concentration-time curves and associated dynamic variables. Exploratory analysis using such kinematic display software, provides a rapid, visually concrete impression of the relationships present in kinetic and dynamic data before the application of standard statistical routines.
The basic idea for the realization of effective statistical data analysis is illustrated with an example. The use of statistical models is explained and the feasibility of objective comparison of the models by an information criterion AIC is demonstrated. Further, the possibility of practical use of Bayesian models for complex data analysis is explained. Finally, the necessity of cooperation between the experts of respective fields and statisticians for further development of statistical data analysis is mentioned.
MOTIVATION: Protein sequence comparison methods are routinely used to infer the intricate network of evolutionary relationships found within the rapidly growing library of protein sequences, and thereby to predict the structure and function of uncharacterized proteins. In the present study, we detail an improved statistical benchmark of pairwise protein sequence comparison algorithms. We use bootstrap resampling techniques to determine standard statistical errors and to estimate the confidence of our conclusions. We show that the underlying structure within benchmark databases causes Efron's standard, non-parametric bootstrap to be biased. Consequently, the standard bootstrap underpredicts average performance when used in the context of evaluating sequence comparison methods. We have developed, as an alternative, an unbiased statistical evaluation based on the Bayesian bootstrap, a resampling method operationally similar to the standard bootstrap. RESULTS: We apply our analysis to the comparative study of amino acid substitution matrix families and find that using modern matrices results in a small, but statistically significant improvement in remote homology detection compared with the classic PAM and BLOSUM matrices. AVAILABILITY: The sequence sets and code for performing these analyses are available from http://compbio.berkeley.edu/. CONTACT: brenner@compbio.berkeley.edu.
Many studies show strong variation of health consumption between regions, suggesting that theses variations are related to the uncertainty of medical practice or to other factors related to health services or patients attitude. However the statistical interpretation of these variations is far from easy: apart from usual and specific information bias, there are statistical problems when observing incidence of events like health care consumption: it is in fact a rare event, which is observed within small population, and among regions with unequal number of person. Therefore, most of the variation reported might be well explained by a purely statistical phenomenon. This paper presents some aspects of this variability for three common indicators of variation, and suggest the use of ad hoc simulation to get statistical criteria.
Similar conserved structures appear in apparently unrelated protein families. Thus, the superfamily of insulin shows an evolutionary relationship with the alpha-conotoxins of marine fish-hunting snails as indicated by methods of protein comparison. In order to reach statistical significance, the A-chains of different insulins, insulin-like growth factors, relaxins, insulin related peptides from invertebrates were drawn for comparison. These data were correlated with sequences from randomly chosen proteins. The alpha-conotoxins show identity scores up to 37.5% and similarity up to 56.2% toward the members of the insulin-superfamily. These scores conform to values achieved by comparing the relaxin and the insulin/IGF-sequences. The data show clearly that the identity and similarity values obtained in the comparison with the insulins are significantly higher than the scores of randomly chosen protein primary structures. According to our calculated data, this hormone system regulating metabolism and growth in vertebrates and the mentioned toxin-receptor system share the same evolutionary ancestor. However, this statistical approach has to be substantiated on gene level.
Gene finding is complicated in organisms that exhibit insertional RNA editing. Here, we demonstrate how our new algorithm Predictor of Insertional Editing (PIE) can be used to locate genes whose mRNAs are subjected to multiple frameshifting events, and extend the algorithm to include probabilistic predictions for sites of nucleotide insertion; this feature is particularly useful when designing primers for sequencing edited RNAs. Applying this algorithm, we successfully identified the nad2, nad4L, nad6 and atp8 genes within the mitochondrial genome of Physarum polycephalum, which had gone undetected by existing programs. Characterization of their mRNA products led to the unanticipated discovery of nucleotide deletion editing in Physarum. The deletion event, which results in the removal of three adjacent A residues, was confirmed by primer extension sequencing of total RNA. This finding is remarkable in that it comprises the first known instance of nucleotide deletion in this organelle, to be contrasted with nearly 500 sites of single and dinucleotide addition in characterized mitochondrial RNAs. Statistical analysis of this larger pool of editing sites indicates that there are significant biases in the 2 nt immediately upstream of editing sites, including a reduced incidence of nucleotide repeats, in addition to the previously identified purine-U bias.
The P-value is the significance probability of obtaining a value of the test statistic that is as extreme, in relation to the null hypothesis, as that observed. Medical researchers may, in some situations, disagree on its appropriate use or on its interpretation as a summary measure of consistency with the null hypothesis in a particular data set. More informative statistical measures such as the likelihood ratio and the Bayesian posterior probability have been suggested for drawing inferences from clinical trials and epidemiologic studies. Causal inference is not statistical in nature; rather it strives to provide scientific explanations or criticisms of proposed explanations that would describe the observed data pattern. In this context, it is important to remember that a finding may not be medically important, or a causal hypothesis may even not be true even if a study shows a significant P-value.
The root mean square successive difference (RMSSD) in heart period series is a time domain measure of heart period variability. The RMSSD is sensitive to high-frequency heart period fluctuations in the respiratory frequency range and has been used as an index of vagal cardiac control. By transfer function simulations, the RMSSD statistic is shown to represent a high-pass filter that effectively captures respiratory sinus arrhythmia but also passes lower frequency fluctuations that can include sympathetic influences. These simulations, together with analysis of actual heart period series, reveal that the RMSSD is biased by basal heart period. Although between-subjects levels of RMSSD covary highly with spectral estimates of high-frequency variability, within-subject RMSSD change scores account for only 50-60% of the variance in spectral estimates. The present findings raise caveats in the applications and interpretation of the RMSSD statistic.
The increased use of effect sizes in single studies and meta-analyses raises new questions about statistical inference. Choice of an effect-size index can have a substantial impact on the interpretation of findings. The authors demonstrate the issue by focusing on two popular effect-size measures, the correlation coefficient and the standardized mean difference (e.g., Cohen's d or Hedges's g), both of which can be used when one variable is dichotomous and the other is quantitative. Although the indices are often practically interchangeable, differences in sensitivity to the base rate or variance of the dichotomous variable can alter conclusions about the magnitude of an effect depending on which statistic is used. Because neither statistic is universally superior, researchers should explicitly consider the importance of base rates to formulate correct inferences and justify the selection of a primary effect-size statistic.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The recent advances in the prediction of intrinsically disordered proteins and the use of protein disorder prediction in the fields of molecular biology and bioinformatics are reviewed here, especially with regard to protein function. First, a close look is taken at intrinsically disordered proteins and then at the methods used for their experimental characterization. Next, the major statistical properties of disordered regions are summarized, and prediction models developed thus far are described, including their numerous applications in functional proteomics. The future of the prediction of protein disorder and the future uses of such predictions in functional proteomics comprise the last section of this article.