PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Interpretation, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Random assignment of available cases: bootstrap standard errors and confidence intervals.

A frequently used experimental design in psychological research randomly divides a set of available cases, a local population, between 2 treatments and then applies an independent-samples t test to either test a hypothesis about or estimate a confidence interval (CI) for the population mean difference in treatment response. C. S. Reichardt and H. F. Gollob (1999) established that the t test can be conservative for this design-yielding hypothesis test P values that are too large or CIs that are too wide for the relevant local population. This article develops a less conservative approach to local population inference, one based on the logic of B. Efron's (1979) nonparametric bootstrap. The resulting randomization bootstrap is then compared with an established approach to local population inference, that based on randomization or permutation tests. Finally, the importance of local population inference is established by reference to the distinction between statistical and scientific inference.

Confidence Intervals↗

Use and misuse of p-values in designed and observational studies: guide for researchers and reviewers.

Analysis of scientific data involves many components, one of which is often statistical testing with the calculation of p-values. However, researchers too often pepper their papers with p-values in the absence of critical thinking about their results. In fact, statistical tests in their various forms address just one question: does an observed difference exceed that which might reasonably be expected solely as a result of sampling error and/or random allocation of experimental material? Such tests are best applied to the results of designed studies with reasonable control of experimental error and sampling error, as well as acquisition of a sufficient sample size. Nevertheless, attributing an observed difference to a specific treatment effect requires critical thinking on the part of the scientist. Observational studies involve data sets whose size is usually a matter of convenience with results that reflect a number of potentially confounding factors. In this situation, statistical testing is not appropriate and p-values may be misleading; other more modern statistical tools should be used instead, including graphic analysis, computer-intensive methods, regression trees, and other procedures broadly classified as bioinformatics, data mining, and exploratory data analysis. In this review, the utility of p-values calculated from designed experiments and observational studies are discussed, leading to the formation of a decision tree to aid researchers and reviewers in understanding both the benefits and limitations of statistical testing.

Clinical Trials as Topic↗

[Analysis of correlated data in occupational medicine: examples with binary data].

In a previous paper in this Journal we presented and discussed examples of analysis of correlated data when the response variable was continuous and normally distributed (the measurement of exposure to a toxic substance was the case in point). In this paper we extend the analysis and the discussion to take into account categorical (binary) variables (described in terms of proportions or odds); to favour the comprehension of the analogies (and discrepancies) between the two contexts we have fully developed an example that mimics the situation presented in the previous paper. Marginal, conditional, random effects and transitional models for correlated data are introduced in practical terms; the meaning of the different estimates obtained are interpreted for epidemiological purposes; the disadvantages of not considering correlation in the analysis are explained and the complexities connected to this type of analysis are fully appreciated. It is concluded that correlated data are very frequently encountered in occupational settings and that an appropriate analysis is necessary. This analysis requires sophisticated computer programs and statistical expertise, particularly in the case of categorical data.

Cluster Analysis↗

Statistics: can we prove an association for a rare complication?

BACKGROUND AND OBJECTIVES: Microcatheters for continuous spinal anesthesia were withdrawn from the market after an apparent increase in the incidence of cauda equina syndrome (CES) associated with spinal anesthesia after introduction of these catheters. The objective of this review is to evaluate the historical data on CES after spinal anesthesia and to compare the result to the recent data. METHODS: The literature on the use of statistics and on complications associated with spinal anesthesia was reviewed. Poisson probabilities were calculated to assess the probability of seeing the recent cases of CES, given the historical data. The sample size required for a prospective study was calculated. RESULTS: Statistics cannot "prove" an hypothesis but can only support or fail to support it. Use and interpretation of the p value are discussed, as are possible problems with interpretation of the p value. The difference between causal and statistical inference is discussed. Probabilities for the occurrence of CES after microcatheter use were calculated. The sample size for a prospective study of this problem was calculated; a large sample is required. CONCLUSIONS: Statistics alone cannot support an association of microcatheters with CES after spinal anesthesia. Additional considerations suggest a possible association, but further study is required.

Anesthesia, Spinal↗

Concurrent validation of substance abusers self-reports against collateral information: percentage agreement vs. kappa vs. Yule's Y.

The ability for chemical users to give an accurate self-report of substance use vs. abstinence has been questioned. This study investigated its concurrent validity, against collateral ratings. The results indicated that validity of reports of chemical use must be evaluated in the context of the validity of other types of information. Chemical use items were corroborated about as often as such nonchemical use items as reports of emergency room visits, arrests, and hospitalizations, thus arguing against the presence of a specific denial syndrome or overarching tendency toward self-misrepresentation. Relative concurrent validities seemed more a function of such factors as item salience and specificity. No consistent trend in patient over- or underreporting of chemical use was found. The choice of concurrent validation statistic is important and can influence interpretation of results. Current standards such as percentage agreement and kappa were seen as flawed; comparisons of results based on these two measures, as well as Yule's Y led to the conclusion that Yule's Y is the statistic of choice.

Adult↗

The kappa statistic in rehabilitation research: an examination.

The number and sophistication of statistical procedures reported in medical rehabilitation research is increasing. Application of the principles and methods associated with evidence-based practice has contributed to the need for rehabilitation practitioners to understand quantitative methods in published articles. Outcomes measurement and determination of reliability are areas that have experienced rapid change during the past decade. In this study, distinctions between reliability and agreement are examined. Information is presented on analytical approaches for addressing reliability and agreement with the focus on the application of the kappa statistic. The following assumptions are discussed: (1) kappa should be used with data measured on a categorical scale, (2) the patients or objects categorized should be independent, and (3) the observers or raters must make their measurement decisions and judgments independently. Several issues related to using kappa in measurement studies are described, including use of weighted kappa, methods of reporting kappa, the effect of bias and prevalence on kappa, and sample size and power requirements for kappa. The kappa statistic is useful for assessing agreement among raters, and it is being used more frequently in rehabilitation research. Correct interpretation of the kappa statistic depends on meeting the required assumptions and accurate reporting.

Activities of Daily Living↗

Statistical intelligence: effective analysis of high-density microarray data.

Microarrays enable researchers to interrogate thousands of genes simultaneously. A crucial step in data analysis is the selection of subsets of interesting genes from the initial set of genes. In many cases, especially when comparing genes expressed in a specific condition to a reference condition, the genes of interest are those which are differentially regulated. This review focuses on the methods currently available for the selection of such genes. Fold change, unusual ratio, univariate testing with correction for multiple experiments, ANOVA and noise sampling methods are reviewed and compared.

Base Sequence↗

A mixture model approach to detecting differentially expressed genes with microarray data.

An exciting biological advancement over the past few years is the use of microarray technologies to measure simultaneously the expression levels of thousands of genes. The bottleneck now is how to extract useful information from the resulting large amounts of data. An important and common task in analyzing microarray data is to identify genes with altered expression under two experimental conditions. We propose a nonparametric statistical approach, called the mixture model method (MMM), to handle the problem when there are a small number of replicates under each experimental condition. Specifically, we propose estimating the distributions of a t -type test statistic and its null statistic using finite normal mixture models. A comparison of these two distributions by means of a likelihood ratio test, or simply using the tail distribution of the null statistic, can identify genes with significantly changed expression. Several methods are proposed to effectively control the false positives. The methodology is applied to a data set containing expression levels of 1,176 genes of rats with and without pneumococcal middle ear infection.

Animals↗

Confidence limit analyses should replace power calculations in the interpretation of epidemiologic studies.

Frequently, after an epidemiologic study is completed, statistical power to detect a relative risk of interest is recalculated using data obtained during the course of the study. A negative study may then be dismissed on the grounds that its power was too low. However, post hoc power calculations ignore the actual relative estimate and its variance, which are by then known. We present evidence that post-study power calculations have little value and should be replaced by a more informative method using the upper (1 - alpha)% confidence limit of the point estimate that touches the value of the relative risk of interest.

Confidence Intervals↗

Fitting and interpreting loglinear interactions in cross-classifications from health policy and medicine.

A nontechnical exposition is presented of current statistical techniques for the analysis of multidimensional tables of counted data. Performing an original analysis of a data set of interest to researchers in health policy and medicine, the paper considers what kinds of questions an analysis by loglinear modeling can address, and what kinds of answers it can obtain and how they may be sought. Unlike most previous expository accounts seeking to provide introductions to this field, this paper does not require a background from the reader in either regression or the analysis of variance. By a thoroughgoing use of odds ratios and higher-order odds ratios, it nevertheless provides a technically accurate account of the key concepts of higher-order interactions among variables, and of models being hierarchical. Statistically more advanced readers are provided with a means of effectively expositing their loglinear modeling methods and conclusions to nonstatisticians; a number of footnotes are directed toward such readers.

California↗

Dichotomizing continuous outcome variables: dependence of the magnitude of association and statistical power on the cutpoint.

Dichotomizing a continuous outcome variable casts that variable in traditional epidemiologic terms (that is, disease, no disease). One consequence is overall reduced statistical power. A more fundamental concern is that the magnitude of various measures of association (for example, prevalence ratio, odds ratio) and statistical power depend on the cutpoint used to dichotomize the variable. The phenomenon is illustrated with a hypothetical situation assuming a two-level predictor variable and a normally distributed outcome variable. As the cutpoint is increased from lower to higher values, the prevalence ratio increases steadily, the odds ratio is described by a U-shaped curve, and statistical power is described by an inverted U-shaped curve. Furthermore, the extent of these effects depends on the difference between the means of the continuous outcome variable for the two levels of the predictor variable. An empirical example is given using data on education and blood pressure (dichotomized to create a high blood pressure vs low blood pressure variable). Except at each end of the distribution, the results follow the hypothetical example. The observation has implications for public health and medical treatment; different cutpoints should be examined to determine the optimal cutpoint in terms of policy and/or treatment decisions. The observation described here also has implications for statistical interpretation; statements about the magnitude of association or statistical significance have limited meaning unless both the cutpoint and the distribution of the outcome variable are specified.

Bias↗

[Cohen's kappa or McNemar's test? A comparison of binary repeated measurements].

This article intends to illustrate the combination of McNemar's significance test and Cohen's kappa coefficient in the comparison of repeated binary measurements. Both methods are standard statistical tools of major relevance for the evaluation and comparison of clinical imaging methods and thus have an impact on the corresponding publications. The interpretation of results obtainable with these methods will be illustrated to facilitate their use based on recent statistical software. Examples will further outline limitations and possible pitfalls in their application to clinical data.

Data Interpretation, Statistical↗

Singular value decomposition analysis of protein sequence alignment score data.

One of the standard tools for the analysis of data arranged in matrix form is singular value decomposition (SVD). Few applications to genomic data have been reported to date mainly for the analysis of gene expression microarray data. We review SVD properties, examine mathematical terms and assumptions implicit in the SVD formalism, and show that SVD can be applied to the analysis of matrices representing pairwise alignment scores between large sets of protein sequences. In particular, we illustrate SVD capabilities for data dimension reduction and for clustering protein sequences. A comparison is performed between SVD-generated clusters of proteins and annotation reported in the SWISS-PROT Database for a set of protein sequences forming the calycin superfamily, entailing all entries corresponding to the lipocalin, cytosolic fatty acid-binding protein, and avidin-streptavidin Prosite patterns.

Amino Acid Sequence↗

Sequence analysis and population data of short tandem repeat polymorphisms at loci D8S639 and D11S488.

Short tandem repeat loci are ideal markers for forensic and paternity case work. A high degree of polymorphism, as determined by gross length measurement, is very often due to complex underlying sequence variation. In the present study, we have studied the sequence structure and population genetics of two short tandem repeat polymorphisms at loci D8S639 and D11S488 in German Caucasians from the region of Hesse. Sequence data revealed a considerable polymorphism at both loci. Locus D8S639 is characterized by a tetranucleotide (AGAT)n repeat pattern with (GAT) and (AGGT) repeats dispersed throughout several alleles. These microvariations lead to alleles differing by one base pair or alleles of identical size. At locus D8S639 we observed 17 allelic lengths comprising 25 different alleles. Alleles at locus D11S488 possess a compound repeat region consisting of (AAAG)n and (GAAG)n repeats. At locus D11S488 we observed 15 allelic lengths with a total of 24 alleles. Allelic lengths increased in size by 4bp increments corresponding to the addition of one tetranucleotide repeat unit. Population data of loci D8S639 and D11S488 revealed a high polymorphism with heterozygosity rates of 0.85 (D8S639) and 0.91 (D11S488).

Alleles↗

Gene assignment by polymerase chain reaction: localization of the human potassium channel IsK gene to the Down's syndrome region of chromosome 21q22.1-q22.2.

Gene mapping, using the polymerase chain reaction (PCR) on DNA obtained from a human/rodent hybrid cell line carrying only the human chromosome 21, permitted the assignment of the human IsK gene, encoding a slowly activating potassium channel, to chromosome 21. PCR analysis of two complete panels of human/rodent hybrid DNA mapped IsK to chromosome 21 with 100% concordance. By performing PCR on DNA of a human chromosome 21 regional mapping panel the gene was sublocalized to chromosome 21q22.1-q22.2, which also contains the putative Down's syndrome (trisomy 21) region. The PCR product obtained from the hybrid cell line DNA carrying only human chromosome 21 was sequenced, thus confirming that the PCR product was derived from human IsK.

Base Sequence↗

Review of Survware from CompuStat software.

Survware is a simple program which permits the user to analyze the responses from multiple choice questionnaires and surveys of 2-50 questions. It is inexpensive when compared with a complete statistical package, but the limitation to frequency data may result in the need to purchase additional software. This program is most likely to be useful to users who regularly collect survey data and who do not need inferential statistical analysis of those data.

Data Collection↗