PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Hypothesis testing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Introduction to biostatistics: Part 4, statistical inference techniques in hypothesis testing.

Statistical methods used to test the null hypothesis are termed tests of significance. Selection of an appropriate test of significance is dependent on the type of data to be analyzed and the number of groups to be compared. Parametric tests of significance are based on the parameters, mean, standard deviation, and variance, and thus are used appropriately when interval or ratio data are analyzed. The t-test and analysis of variance (ANOVA) are examples of parametric tests of significance. Assumptions regarding the data to be analyzed when using the t-test or ANOVA include normality of the populations from which the sample data are drawn, homogeneity of the variances of the populations from which the sample data are drawn, and independence of the data points within a sample group. The t-test is the appropriate test of significance to use if there are only two groups to compare. If there are three or more groups to compare, ANOVA is the appropriate test. ANOVA holds the preset alpha level constant. While ANOVA will imply a significant difference between the groups compared, a multiple comparison test will define which of the three or more groups differ significantly.

Analysis of Variance↗

Hypothesis testing approaches to the exon prediction problem.

MOTIVATION: Many gene identification methods assign scores to gene elements prior to their assembly into predicted genes. The scoring system is often based on log-likelihood ratios. These methods usually perform well but it is difficult to interpret how significant a score is. RESULTS: We have developed several tests of significance for the scores: (1) a sum-of-scores test (SST), (2) an intersection-union test (IUT), based on a multiple hypothesis testing interpretation of an exon's score and (3) a meta-analytical approach (MA), which combines several P-values, corresponding to the exon's parts, to yield a global P-value. We performed simulation studies, which show that the MA has better sensitivity and specificity than other methods and is easier to interpret by non-expert users. This is an improvement over other methods and is especially relevant for users who would like to predict incomplete gene sequences.

Algorithms↗

Information Selection and Utilization in Hypothesis Testing: A Comparison of Process-Tracing and Structural Analysis Techniques.

This paper examines the contributions of process-tracing and structural analysis techniques to our understanding of the cognitive processes underlying hypothesis testing. An experiment designed to compare and contrast the techniques is reported in which participants (N = 72) made diagnostic judgments by selecting information they felt was necessary from computerized information-boards (Mouselab 4.2: Johnson et al., 1989). This innovative methodology enabled us to explore the potential of a multimethod approach in providing insights into how people's understanding of the diagnostic value of conditional probability information in hypothesis testing deviated from the prescriptions of Bayes' theorem. In short, the process-tracing results indicated that, as well as typically selecting information according to a strategy which did not fit the Bayesian prediction, participants generally paid most attention to information directly pertinent to the hypothesis or hypotheses they were explicitly asked to evaluate. Although the structural analysis also implied that participants' integration strategies were not strictly Bayesian, direct comparisons between the data sources highlighted possible dangers associated with relying on each technique in isolation. Most significantly, by synthesizing data between techniques the value of adopting combined process-tracing/structural analysis approaches in decision research was demonstrated. Copyright 1998 Academic Press.

Journal Article↗

Hypothesis testing of a change point during cognitive decline among Alzheimer's disease patients.

In this paper, we present a statistical hypothesis test for detecting a change point over the course of cognitive decline among Alzheimer's disease patients. The model under the null hypothesis assumes a constant rate of cognitive decline over time and the model under the alternative hypothesis is a general bilinear model with an unknown change point. When the change point is unknown, however, the null distribution of the test statistics is not analytically tractable and has to be simulated by parametric bootstrap. When the alternative hypothesis that a change point exists is accepted, we propose an estimate of its location based on the Akaike's Information Criterion. We applied our method to a data set from the Neuropsychological Database Initiative by implementing our hypothesis testing method to analyze Mini Mental Status Exam scores based on a random-slope and random-intercept model with a bilinear fixed effect. Our result shows that despite large amount of missing data, accelerated decline did occur for MMSE among AD patients. Our finding supports the clinical belief of the existence of a change point during cognitive decline among AD patients and suggests the use of change point models for the longitudinal modeling of cognitive decline in AD research.

Aged↗

Interval estimation and hypothesis testing of intraclass correlation coefficients: the generalized variable approach.

In this paper, we propose a novel approach using the concept of generalized variable (GV) for the confidence interval estimation of the difference of two intraclass correlation coefficients under unequal family sizes. This approach can also easily provide P-values for hypothesis testing. Simulation results show that the GV approach can provide confidence intervals with good coverage properties and perform hypothesis testing with satisfactory type-I error control. Furthermore, the confidence intervals and P-values by GV approach can be easily obtained by simulation. Therefore the GV approach is a suitable candidate for making inference concerning two intraclass correlation coefficients.

Biometry↗

Microbial biodiversity: approaches to experimental design and hypothesis testing in primary scientific literature from 1975 to 1999.

Research interest in microbial biodiversity over the past 25 years has increased markedly as microbiologists have become interested in the significance of biodiversity for ecological processes and as the industrial, medical, and agricultural applications of this diversity have evolved. One major challenge for studies of microbial habitats is how to account for the diversity of extremely large and heterogeneous populations with samples that represent only a very small fraction of these populations. This review presents an analysis of the way in which the field of microbial biodiversity has exploited sampling, experimental design, and the process of hypothesis testing to meet this challenge. This review is based on a systematic analysis of 753 publications randomly sampled from the primary scientific literature from 1975 to 1999 concerning the microbial biodiversity of eight habitats related to water, soil, plants, and food. These publications illustrate a dominant and growing interest in questions concerning the effect of specific environmental factors on microbial biodiversity, the spatial and temporal heterogeneity of this biodiversity, and quantitative measures of population structure for most of the habitats covered here. Nevertheless, our analysis reveals that descriptions of sampling strategies or other information concerning the representativeness of the sample are often missing from publications, that there is very limited use of statistical tests of hypotheses, and that only a very few publications report the results of multiple independent tests of hypotheses. Examples are cited of different approaches and constraints to experimental design and hypothesis testing in studies of microbial biodiversity. To prompt a more rigorous approach to unambiguous evaluation of the impact of microbial biodiversity on ecological processes, we present guidelines for reporting information about experimental design, sampling strategies, and analyses of results in publications concerning microbial biodiversity.

Bacteria↗

An interactive power analysis tool for microarray hypothesis testing and generation.

MOTIVATION: Human clinical projects typically require a priori statistical power analyses. Towards this end, we sought to build a flexible and interactive power analysis tool for microarray studies integrated into our public domain HCE 3.5 software package. We then sought to determine if probe set algorithms or organism type strongly influenced power analysis results. RESULTS: The HCE 3.5 power analysis tool was designed to import any pre-existing Affymetrix microarray project, and interactively test the effects of user-defined definitions of alpha (significance), beta (1-power), sample size and effect size. The tool generates a filter for all probe sets or more focused ontology-based subsets, with or without noise filters that can be used to limit analyses of a future project to appropriately powered probe sets. We studied projects from three organisms (Arabidopsis, rat, human), and three probe set algorithms (MAS5.0, RMA, dChip PM/MM). We found large differences in power results based on probe set algorithm selection and noise filters. RMA provided high sensitivity for low numbers of arrays, but this came at a cost of high false positive results (24% false positive in the human project studied). Our data suggest that a priori power calculations are important for both experimental design in hypothesis testing and hypothesis generation, as well as for the selection of optimized data analysis parameters. AVAILABILITY: The Hierarchical Clustering Explorer 3.5 with the interactive power analysis functions is available at www.cs.umd.edu/hcil/hce or www.cnmcresearch.org/bioinformatics. CONTACT: jseo@cnmcresearch.org

Algorithms↗

Hypothesis testing I: proportions.

Statistical inference involves two analysis methods: estimation and hypothesis testing, the latter of which is the subject of this article. Specifically, Z tests of proportion are highlighted and illustrated with imaging data from two previously published clinical studies. First, to evaluate the relationship between nonenhanced computed tomographic (CT) findings and clinical outcome, the authors demonstrate the use of the one-sample Z test in a retrospective study performed with patients who had ureteral calculi. Second, the authors use the two-sample Z test to differentiate between primary and metastatic ovarian neoplasms in the diagnosis and staging of ovarian cancer. These data are based on a subset of cases from a multiinstitutional ovarian cancer trial conducted by the Radiologic Diagnostic Oncology Group, in which the roles of CT, magnetic resonance imaging, and ultrasonography (US) were evaluated. The statistical formulas used for these analyses are explained and demonstrated. These methods may enable systematic analysis of proportions and may be applied to many other radiologic investigations.

Diagnostic Imaging↗

The interaction of implicit learning, explicit hypothesis testing learning and implicit-to-explicit knowledge extraction.

To further explore the interaction between the implicit and explicit learning processes in skill acquisition (which have been tackled before, e.g. in [Sun, R., Merrill, E., & Peterson, T. (2001). From implicit skill to explicit knowledge: A bottom-up model of skill learning. Cognitive Science, 25(2), 203-244; Sun, R., Slusarz, P., & Terry, C. (2005). The interaction of the explicit and the implicit in skill learning: A dual-process approach. Psychological Review, 112(1), 159-192]), this paper explores details of the interaction of different learning modes: implicit learning, explicit hypothesis testing learning, and implicit-to-explicit knowledge extraction. Contrary to the common tendency in the literature to study each type of learning in isolation, this paper highlights the interaction among them and various effects of the interaction on learning, including the synergy effect. This work advocates an integrated model of skill learning that takes into account both implicit and explicit learning processes; moreover, it also uniquely embodies a bottom-up (implicit-to-explicit) learning approach in addition to other types of learning. The paper shows that this model accounts for various effects in the human behavioural data from the psychological experiments with the process control task, in addition to accounting for other data in other psychological experiments (which has been reported elsewhere). The paper shows that to account for these effects, implicit learning, bottom-up implicit-to-explicit extraction and explicit hypothesis testing learning are all needed.

Algorithms↗

Mining of biological data I: identifying discriminating features via mean hypothesis testing.

Large volumes of data are routinely collected during bioprocess operations and, more recently, in basic biological research using genomics-based technologies. While these data often lack sufficient detail to be used for mechanism identification, it is possible that the underlying mechanisms affecting cell phenotype or process outcome are reflected as specific patterns in the overall or temporal sensor logs. This raises the possibility of identifying outcome-specific fingerprints that can be used for process or phenotype classification and the identification of discriminating characteristics, such as specific genes or process variables. The aim of this work is to provide a systematic approach to identifying and modeling patterns in historical records and using this information for process classification. This approach differs from others in that emphasis is placed on analyzing the data structure first and thereby extracting potentially relevant features prior to model creation. The initial step in this overall approach is to first identify the discriminating features of the relevant measurements and time windows, which can then be subsequently used to discriminate among different classes of process behavior. This is achieved via a mean hypothesis testing algorithm. Next, the homogeneity of the multivariate data in each class is explored via a novel cluster analysis technique called PC1 Time Series Clustering to ensure that the data subsets used accurately reflect the variability displayed in the historical records. This will be the topic of the second paper in this series. We present here the method for identifying discriminating features in data via mean hypothesis testing along with results from the analysis of case studies from industrial fermentations

Algorithms↗

Two-dimensional protein electrophoresis and multiple hypothesis testing to detect potential serum protein biomarkers in children with fetal alcohol syndrome.

Fetal alcohol syndrome (FAS) surveillance and intervention efforts are hampered by the lack of a specific biochemical test for diagnosis of the syndrome. Based on the hypothesis that abnormalities in growth and development (key features of FAS) involve altered protein metabolism, we analyzed serum proteins by two-dimensional gel electrophoresis and image analysis to search for potential protein biomarkers of FAS. Serum samples from 12 participants in whom FAS had been diagnosed and 8 sex- and age-matched participants whose mothers did not consume alcohol were analyzed in duplicate to determine whether the integrated intensities of matched proteins are significantly altered in children with FAS. Multiple hypothesis testing on 34 of the gels consisting of more than 1700 spots per gel revealed 21 proteins that we classified as potential protein biomarkers of FAS on the basis of significant t-test differences at p < 0.02. We classified 8 of the proteins as candidate biomarkers on the basis of significant concentration differences between case and control subjects at p < 0.01. One of the proteins is clearly an isoform of retinol binding protein; two appear in the area of the gel where alcohol dehydrogenase is expected to appear; one appears to be an isoform of alpha-1-antitrypsin; three appear to be isoforms of the beta-chain of haptoglobin; three may be forms of immunoglobulin light chains; and several others have not been associated with known proteins. No single protein differentiated all case subjects from control subjects, but stepwise canonical discriminant analyses revealed four groups of spots that distinguished between FAS case and control subjects with no misclassifications.(ABSTRACT TRUNCATED AT 250 WORDS)

Biomarkers↗

Hypothesis testing of time-dependent recurrent events.

This paper examines the problem of hypothesis testing for comparison of time-dependent recurrent events between categorical exposure groups. We compare three methods (a summary chi-square test based on a generalization of the simple Poisson distribution, a chi-square test based on a generalization of the compound Poisson distribution, and a test on risk scores based on individual observed to expected ratios), when the dependent variable may be autocorrelated within an individual, and disease risk may be heterogeneous among subjects. We present a simulation study and an application to a cohort of sickle cell anaemia patients. With autocorrelation or heterogeneous risk present, the simple chi-square test is inappropriate, while the other two methods perform well. An attraction of the risk score method is its relative ease of application.

Anemia, Sickle Cell↗

Hypothesis testing in patients with persecutory delusions: comparison with depressed and normal subjects.

The hypothesis-testing skills of patients with persecutory delusions were studied, and compared with those of matched depressed and normal control groups. Subjects were required to complete a series of visual discrimination problems in which they had to choose between pairs of stimuli presented on cards. Following positive or negative feedback from the examiner, subjects' ability to progressively narrow down the set of possible correct solutions was assessed. The groups did not differ in the range or total number of hypotheses generated. The deluded subjects were less inclined than the controls to stick to their hypotheses when given positive feedback and were more inclined to stick to their hypotheses following negative feedback. They also showed less evidence of 'focusing' down their hypothesis to an overall correct solution, in response to successive feedback.

Adult↗

Hypothesis testing for data from different family study designs.

A test statistic that is valid for data collected according to a particular type of family study design is not necessarily valid when applied to data obtained from a different type of family study design. When this can occur, a different test that usually is valid is developed for each type of family study design. However, investigators might find that their data come from two (or more) different family study designs, each requiring a different test, yet they want an overall conclusion, essentially a valid hypothesis test that is as powerful as possible. When the underlying genetic model is unknown, it is not clear how to proceed, as several alternative approaches might appear feasible. By using as an example the development of a test of association for data concerning affected singletons and their parents and affected sib pairs and their parents, it is shown that it may not be possible to develop a universally optimal approach without knowledge of the underlying genetic model.

Data Interpretation, Statistical↗

Statistical sampling and hypothesis testing in orthopaedic research.

The purpose of the current article was to review the process of hypothesis testing and statistical sampling and empower readers to critically appraise the literature. When the p value of a study lies above the alpha threshold, the results are said to be not statistically significant. It is possible, however, that real differences do exist, but the study was insufficiently powerful to detect them. In that case, the conclusion that two groups are equivalent is wrong. The probability of this mistake, the Type II error, is given by the beta statistic. The complement of beta, or 1-beta, representing the chance of avoiding a Type II error, is termed the statistical power of the study. We previously examined the statistical power and sample size in all of the studies published in 1997 in the American and British volumes of the Journal of Bone and Joint Surgery, and in Clinical Orthopaedics and Related Research. In the journals examined, only 3% of studies had adequate statistical power to detect a small effect size in this sample. In addition, a study examining only randomized control trials in these journals showed that none of 25 randomized control trials had adequate statistical power to detect a small effect size. However, beta, or power, is less well understood. Because of this, researchers and readers should be aware of the need to address issues of statistical power before a study begins and be cautious of studies that conclude that no difference exists between groups.

Epidemiologic Research Design↗

Hypothesis testing under mixture models: application to genetic linkage analysis.

In this paper we propose a new class of statistics to test a simple hypothesis against a family of alternatives characterized by a mixture model. Unlike the likelihood ratio statistic, whose large sample distribution is still unknown in this situation, these new statistics have a simple asymptotic distribution to which to refer under the null hypothesis. Simulation results suggest that it has adequate power in detecting the alternatives. Its application to genetic linkage analysis in the presence of the genetic heterogeneity that motivated this work is emphasized.

Biometry↗

Are effect sizes and confidence levels problems for or solutions to the null hypothesis test?

Some have proposed that the null hypothesis significance test, as usually conducted using the t test of the difference between means, is an impediment to progress in psychology. To improve its prospects, using Neyman-Pearson confidence intervals and Cohen's standardized effect sizes, d, is recommended. The purpose of these approaches is to enable us to understand what can appropriately be said about the distances between the means and their reliability. Others have written extensively that these recommended strategies are highly interrelated and use identical information. This essay was written to remind us that the t test, based on the sample--not the true--standard deviation, does not apply solely to distance between means. The t test pertains to a much more ambiguous specification: the difference between samples, including sampling variations of the standard deviation.

Confidence Intervals↗