PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Statistics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

A comparison between the sciences of epidemiology and statistics based on an examination of epidemiological or statistical studies on diabetes in Japan.

The authors selected 24 original papers which were regarded them as the epidemiological study and the statistical study from their titles, from the end of World War II to 1981. And these papers were selected from 3 medical journals of internal medicine, other medical journals and proceedings of 2 International Conferences (see Table 1), and also were the object of study, namely, theoretical considerations. Besides we classified these 24 papers into 2 sorts; papers for an epidemiological study and a statistical study, and made a comparative study of details of these papers theoretically. As the result we were able to clarify what the authors of 24 papers had considered about the natures of epidemiology and statistics as the science. It was clarified that two sciences, epidemiology and statistics, had been in the general trend without any recognition of the differences between two. And as the conclusion we pointed out that the field of activity of statistics was broader than that of epidemiology, and the nature of statistics as the science might be changeable according to the object, moreover, statistical theory might be a branch of mathematics and so on.

Diabetes Mellitus↗

Issues in biomedical statistics: statistical inference.

The first step in making inferences under the frequentist system of statistical logic is to propose a null hypothesis. An experiment is then performed, or a set of observations made. The resulting data are subjected to statistical analysis to determine whether the null hypothesis should be rejected or not. If it is, then some alternative hypothesis must have been entertained. In biomedical work, the alternative hypothesis should usually be non-specific and it follows that the statistical test of the null hypothesis should be interpreted in a two-sided fashion. The decision to reject or accept statistical null hypotheses, whether on the basis of a P value or confidence intervals, is probabilistic in nature and always attended by the risk of error. It is argued that, in biomedical research, it is the risk of making false-positive statistical inferences (Type I error) that should be most closely controlled. The risks of Type I error cannot be considered in isolation from the model of inference under which the null hypothesis is tested. That which forms the basis for using the classical t, F and X2 tests is the population model, in which the inference is referred to a defined population that has been randomly sampled and which conforms to a specified frequency distribution. Under this model, serious errors in statistical inference can occur if the actual distributions of the populations do not conform to those specified by theory. More importantly, the population model is inappropriate to most biomedical research, in which treatment groups are created by randomization but not by random sampling.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals↗

Statistical representation of high-dimensional deformation fields with application to statistically constrained 3D warping.

This paper proposes a 3D statistical model aiming at effectively capturing statistics of high-dimensional deformation fields and then uses this prior knowledge to constrain 3D image warping. The conventional statistical shape model methods, such as the active shape model (ASM), have been very successful in modeling shape variability. However, their accuracy and effectiveness typically drop dramatically in high-dimensionality problems involving relatively small training datasets, which is customary in 3D and 4D medical imaging applications. The proposed statistical model of deformation (SMD) uses wavelet-based decompositions coupled with PCA in each wavelet band, in order to more accurately estimate the pdf of high-dimensional deformation fields, when a relatively small number of training samples are available. SMD is further used as statistical prior to regularize the deformation field in an SMD-constrained deformable registration framework. As a result, more robust registration results are obtained relative to using generic smoothness constraints on deformation fields, such as Laplacian-based regularization. In experiments, we first illustrate the performance of SMD in representing the variability of deformation fields and then evaluate the performance of the SMD-constrained registration, via comparing a hierarchical volumetric image registration algorithm, HAMMER, with its SMD-constrained version, referred to as SMD+HAMMER. This SMD-constrained deformable registration framework can potentially incorporate various registration algorithms to improve robustness and stability via statistical shape constraints.

Algorithms↗

Error in statistical tests of error in statistical tests.

BACKGROUND: A recent paper found that terminal digits of statistical values in Nature deviated significantly from an equiprobable distribution, indicating errors or inconsistencies in rounding. This finding, as well as the discovery that a large percentage of p values were inconsistent with reported test statistics, led to a great deal of concern in the popular press and scientific community. The findings ultimately led to new guidelines for all Nature Research Journals. METHODS: We checked the statistical analysis behind the original paper's tests of equiprobability. RESULTS: The original paper tested equiprobability with the Kolmogorov-Smirnov test outside its regime of validity. Correct tests find no statistically significant deviations from equiprobability for the statistical values in Nature. CONCLUSION: Statistical tests should be used correctly.

Confidence Intervals↗

Advanced statistics: bootstrapping confidence intervals for statistics with "difficult" distributions.

The use of confidence intervals in reporting results of research has increased dramatically and is now required or highly recommended by editors of many scientific journals. Many resources describe methods for computing confidence intervals for statistics with mathematically simple distributions. Computing confidence intervals for descriptive statistics with distributions that are difficult to represent mathematically is more challenging. The bootstrap is a computationally intensive statistical technique that allows the researcher to make inferences from data without making strong distributional assumptions about the data or the statistic being calculated. This allows the researcher to estimate confidence intervals for statistics that do not have simple sampling distributions (e.g., the median). The purposes of this article are to describe the concept of bootstrapping, to demonstrate how to estimate confidence intervals for the median and the Spearman rank correlation coefficient for non-normally-distributed data from a recent clinical study using two commonly used statistical software packages (SAS and Stata), and to discuss specific limitations of the bootstrap.

Confidence Intervals↗

Statistical inference using the g or K point pattern spatial statistics.

Spatial point pattern analysis provides a statistical method to compare an observed spatial pattern against a hypothesized spatial process model. The G statistic, which considers the distribution of nearest neighbor distances, and the K statistic, which evaluates the distribution of all neighbor distances, are commonly used in such analyses. One method of employing these statistics involves building a simulation envelope from the result of many simulated patterns of the hypothesized model. Specifically, a simulation envelope is created by calculating, at every distance, the minimum and maximum results computed across the simulated patterns. A statistical test is performed by evaluating where the results from an observed pattern fall with respect to the simulation envelope. However, this method, which differs from P. Diggle's suggested approach, is invalid for inference because it violates the assumptions of Monte Carlo methods and results in incorrect type I error rate performance. Similarly, using the simulation envelope to estimate the range of distances over which an observed pattern deviates from the hypothesized model is also suspect. The technical details of why the simulation envelope provides incorrect type I error rate performance are described. A valid test is then proposed, and details about how the number of simulated patterns impacts the statistical significance are explained. Finally, an example of using the proposed test within an exploratory data analysis framework is provided.

Data Interpretation, Statistical↗

Adjusting for population heterogeneity: a framework for characterizing statistical information and developing efficient test statistics.

The focus of this work is the TDT-type and family-based test statistics used for adjusting for potential confounding due to population heterogeneity or misspecified allele frequencies. A variety of heuristics have been used to motivate and derive these statistics, and the statistics have been developed for a variety of analytic goals. There appears to be no general theoretical framework, however, that may be used to evaluate competing approaches. Furthermore, there is no framework to guide the development of efficient TDT-type and family-based methods for analytic goals for which methods have not yet been proposed. The purpose of this paper is to present a theoretical framework that serves both to identify the information which is available to methods that are immune to confounding due to population heterogeneity or misspecified allele frequencies, and to inform the construction of efficient unbiased tests in novel settings. The development relies on the existence of a characterization of the null hypothesis in terms of a completely specified conditional distribution of transmitted genotypes. An important observation is that, with such a characterization, when the conditioning event is unobserved or incomplete, there is statistical information that cannot be exploited by any exact conditional test. The main technical result of this work is an approach to computing test statistics for local alternatives that exploit all of the available statistical information.

Data Interpretation, Statistical↗

Further statistics in dentistry. Part 9: Bayesian statistics.

Statistics can be defined as the methods used to assimilate data, so that guidance can be given, and conclusions drawn, in situations which involve uncertainty. In particular, statistical inference is concerned with drawing conclusions about particular aspects of a population when that population cannot be studied in full. Uncertainty arises here because the totality of the information is not available. Instead, to make inferences about the population, it is necessary to rely on a sample of data which is selected from the population; this sample data may be augmented, in certain circumstances, by auxiliary information which is obtained independently of the sample data. Clearly, uncertainty lies at the heart of statistics and statistical inference. This uncertainty is measured by a probability which therefore forms the crux of statistics and must be properly understood in order to interpret a statistical analysis.

Algorithms↗

Tutorials in clinical research: part VII. Understanding comparative statistics (contrast)--part A: general concepts of statistical significance.

OBJECTIVES/HYPOTHESIS: The present tutorial is the seventh in a series of Tutorials in Clinical Research. The specific purpose of the tutorial (Part A) and its sequel (Part B) is to introduce and explain three commonly used statistical tools for assessing contrast in the comparison between two groups. STUDY DESIGN: Tutorial. METHODS: The authors met weekly for 10 months discussing clinical research studies and the applied statistics. The difficulty was not in the material but in the effort to make the report easy to read and as short as possible. RESULTS: The tutorial is organized into two parts. Part A, which is the present report, focuses on the fundamental concepts of the null hypothesis and comparative statistical significance. The sequel, Part B, discusses the application of three common statistical indexes of contrast, the chi2, Mann-Whitney U, and Student t tests. CONCLUSIONS: Assessing the validity of medical studies requires a working knowledge of research design and statistics; obtaining this knowledge need not be beyond the ability of the busy surgeon. The authors have tried to construct an accurate, easy-to-read, easy-to-apply, basic introduction to comparing two groups. The long-term goal of the present tutorial and others in the series is to facilitate basic understanding of clinical research, thereby stimulating reading of some of the numerous well-written research design and statistical texts. This knowledge may then be applied to the continuing educational review of the literature and the systematic prospective analysis of individual practices.

Bias↗

Statistical methods in epidemiology: I. Statistical errors in hypothesis testing.

PURPOSE: Although scientific journal editors are making use of statisticians in the review process, the quality of statistical reporting in many journals remains poor. In many cases the problem for the scientist would appear to be a lack of understanding of basic statistics. The focus of the scientist is on showing 'p < 0.05', when what is actually required is a statement about effect size and interval estimation. The aim of this paper is to show the inadequacy of reporting of results using p-values alone. This paper is the first in a series detailing common statistical methods, with a view to aiding potential authors in their statistical presentation of data. METHOD: A review of the basic hypothesis test, using examples from the author's own teaching experiences. RESULTS: Type I and type II errors are defined; the problem of multiple comparisons is highlighted; interval estimation is introduced. CONCLUSIONS: The case for considering the p-value as an error probability is made which suggests ways of improving statistical presentation and thus expediting the statistical review process.

Confidence Intervals↗

Improved statistical inference from DNA microarray data using analysis of variance and a Bayesian statistical framework. Analysis of global gene expression in Escherichia coli K12.

We describe statistical methods based on the t test that can be conveniently used on high density array data to test for statistically significant differences between treatments. These t tests employ either the observed variance among replicates within treatments or a Bayesian estimate of the variance among replicates within treatments based on a prior estimate obtained from a local estimate of the standard deviation. The Bayesian prior allows statistical inference to be made from microarray data even when experiments are only replicated at nominal levels. We apply these new statistical tests to a data set that examined differential gene expression patterns in IHF(+) and IHF(-) Escherichia coli cells (Arfin, S. M., Long, A. D., Ito, E. T., Tolleri, L., Riehle, M. M., Paegle, E. S., and Hatfield, G. W. (2000) J. Biol. Chem. 275, 29672-29684). These analyses identify a more biologically reasonable set of candidate genes than those identified using statistical tests not incorporating a Bayesian prior. We also show that statistical tests based on analysis of variance and a Bayesian prior identify genes that are up- or down-regulated following an experimental manipulation more reliably than approaches based only on a t test or fold change. All the described tests are implemented in a simple-to-use web interface called Cyber-T that is located on the University of California at Irvine genomics web site.

Bayes Theorem↗

Can census offices publish statistics for more than one small area geography? An analysis of the differencing problem in statistical disclosure.

"The paper describes a problem faced by National Statistical Offices when publishing the results of decennial censuses for small geographical areas. If they publish statistical tables for two or more sets of areas, users can compare the tables and produce new statistics for the areas formed by differencing, which may have populations below confidentiality thresholds. To investigate the problem, the authors construct a software system and carry out a series of experiments using a large synthetic population base for Yorkshire and Humberside [in England]. The results indicate that publishing statistics for zones close in size to the primary areas is not safe unless the zones have been carefully designed. However, publishing statistics for sufficiently large areas such as 5km grid squares or postal sectors alongside enumeration districts is safe."

Censuses↗

Statistical significance versus clinical relevance. Part I. The essential role of the power of a statistical test.

When comparing two treatment groups, hypothesis testing is widely used. However, clinical trialists should be more interested in statistical methods which elicit the magnitude of the differences between treatment groups, rather than a simple indication of whether or not the differences are statistically significant. Statistical significance does not necessarily imply clinical relevance. If the true difference between two treatment groups is so small that it is clinically irrelevant, a sample size can be found for which this difference is statistically significant. On the other hand, if the difference between treatment groups is statistically non-significant, it may still be clinically important. The limitations of conventional hypothesis testing of equal true means as such are highlighted. The need to control the power of the test--which takes into account the difference in treatment means which is considered important (clinically relevant) by the researcher--is discussed.

Clinical Trials as Topic↗

An evaluation of using a web-based statistics test to teach statistics to post-registration nursing students.

Part of evidence-based practice is an ability to appraise research. Studies have shown that understanding the statistical components of a study is an area some nursing students struggle with. This paper reports the implementation and evaluation of a web-based statistics test to teach statistics to post-registration nursing students. The evaluation used both the data from the web-based statistics test, to measure students' learning, and from a survey completed after the module, to examine students' attitudes. The survey included some qualitative elements. The evaluation demonstrated that this is a valid and acceptable method of improving knowledge and understanding of statistics within this group of students.

Adult↗

Statistical learning in a serial reaction time task: access to separable statistical cues by individual learners.

The ability of adult learners to exploit the joint and conditional probabilities in a serial reaction time task containing both deterministic and probabilistic information was investigated. Learners used the statistical information embedded in a continuous input stream to improve their performance for certain transitions by simultaneously exploiting differences in the predictability of 2 or more underlying statistics. Analysis of individual learners revealed that although most acquired the underlying statistical structure veridically, others used an alternate strategy that was partially predictive of the sequences. The findings show that learners possess a robust learning device well suited to exploiting the relative predictability of more than I source of statistical information at the same time. This work expands on previous studies of statistical learning, as well as studies of artificial grammar learning and implicit sequence learning.

Adult↗