PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Hypothesis testing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Basic statistics for clinicians: 1. Hypothesis testing.

In the first of a series of four articles the authors explain the statistical concepts of hypothesis testing and p values. In many clinical trials investigators test a null hypothesis that there is no difference between a new treatment and a placebo or between two treatments. The result of a single experiment will almost always show some difference between the experimental and the control groups. Is the difference due to chance, or is it large enough to reject the null hypothesis and conclude that there is a true difference in treatment effects? Statistical tests yield a p value: the probability that the experiment would show a difference as great or greater than that observed if the null hypothesis were true. By convention, p values of less than 0.05 are considered statistically significant, and investigators conclude that there is a real difference. However, the smaller the sample size, the greater the chance of erroneously concluding that the experimental treatment does not differ from the control--in statistical terms, the power of the test may be inadequate. Tests of several outcomes from one set of data may lead to an erroneous conclusion that an outcome is significant if the joint probability of the outcomes is not taken into account. Hypothesis testing has limitations, which will be discussed in the next article in the series.

Clinical Trials as Topic↗

The effects of introtacts on hypothesis testing in kindergarten and first-grade children.

Kindergarten and first-grade children were trained against their initial dimensional preference in a 2-dimensional simultaneous discrimination learning task. One third of the children received pretraining in using a sequential hypothesis-testing strategy, one third received pretraining in which they experienced solutions to tasks of the same type, and one third received no pretraining. Half of the children received introtact probes prior to each trial in the criterion task. Introtact probes had no effect on the performance of kindergarten children but facilitated the performance of first-grade children who received pretraining. Performance was generally better in the pretraining conditions than in the control condition and was generally better for first graders than for kindergarten children. Indices of the use of the sequential hypothesis-testing strategy were obtained from the responses to introtact probes. 75% of the first graders who received pretraining in hypothesis testing showed high proficiency in using the strategy, whereas only 38% of the kindergarten children did so. A strong tendency to become fixated on the irrelevant dimension was evident at both age levels.

Child↗

Information selection and use in hypothesis testing: what is a good question, and what is a good answer?

The process of hypothesis testing entails both information selection (asking questions) and information use (drawing inferences from the answers to those questions). We demonstrate that although subjects may be sensitive to diagnosticity in choosing which questions to ask, they are insufficiently sensitive to the fact that different answers to the same question can have very different diagnosticities. This can lead subjects to overestimate or underestimate the information in the answers they receive. This phenomenon is demonstrated in two experiments using different kinds of inferences (category membership of individuals and composition of sampled populations). In combination with certain information-gathering tendencies, demonstrated in a third experiment, insensitivity to answer diagnosticity can contribute to a tendency toward preservation of the initial hypothesis. Results such as these illustrate the importance of viewing hypothesis-testing behavior as an interactive, multistage process that includes selecting questions, interpreting data, and drawing inferences.

Adult↗

Hypothesis testing via integrated computer modeling and digital fluorescence microscopy.

Computational modeling has the potential to add an entirely new approach to hypothesis testing in yeast cell biology. Here, we present a method for seamless integration of computational modeling with quantitative digital fluorescence microscopy. This integration is accomplished by developing computational models based on hypotheses for underlying cellular processes that may give rise to experimentally observed fluorescent protein localization patterns. Simulated fluorescence images are generated from the computational models of underlying cellular processes via a "model-convolution" process. These simulated images can then be directly compared to experimental fluorescence images in order to test the model. This method provides a framework for rigorous hypothesis testing in yeast cell biology via integrated mathematical modeling and digital fluorescence microscopy.

Computational Biology↗

Multiway sequential hypothesis testing for tachyarrhythmia discrimination.

A multiway sequential hypothesis testing (M-SHT) algorithm is proposed for simultaneous discrimination of cardiac tachyarrhythmias--supraventricular tachycardia (SVT) and ventricular tachycardia (VT)--from normal sinus rhythm (NSR). The M-SHT algorithm calculates a likelihood function from atrio-ventricular delay measurements, and compares this function with thresholds derived from specified error probabilities for the arrhythmias to be discriminated. Performance of this algorithm was evaluated on dual channel endocardial electrograms recorded in the cardiac electrophysiology laboratory. Two databases were developed, one for development of the algorithm and another for evaluation. The M-SHT algorithm accurately classified 26 out of 28 NSR (2 misclassified as SVT), 31 out of 31 cases of SVT, and 41 out of 43 VT (2 misclassified as NSR). The average length of time taken for classification of the three rhythms was: 3.6 s for NSR, 5.0 s for SVT, and 1.6 s for VT. Unique features of this algorithm are that acceptable error rates for each arrhythmia are independently specified and accuracy can be traded off for a faster detection time, and vice versa.

Adult↗

Answering two criticisms of hypothesis testing: reply to Serlin.

Two criticisms of hypothesis testing have been repeated for half a century. Leventhal (1999) defended against those criticisms. Serlin (2000) commented on Leventhal's paper and criticized parts of Leventhal's defense. Serlin's comments are discussed and his criticisms answered.

Confidence Intervals↗

GeneTools--application for functional annotation and statistical hypothesis testing.

BACKGROUND: Modern biology has shifted from "one gene" approaches to methods for genomic-scale analysis like microarray technology, which allow simultaneous measurement of thousands of genes. This has created a need for tools facilitating interpretation of biological data in "batch" mode. However, such tools often leave the investigator with large volumes of apparently unorganized information. To meet this interpretation challenge, gene-set, or cluster testing has become a popular analytical tool. Many gene-set testing methods and software packages are now available, most of which use a variety of statistical tests to assess the genes in a set for biological information. However, the field is still evolving, and there is a great need for "integrated" solutions. RESULTS: GeneTools is a web-service providing access to a database that brings together information from a broad range of resources. The annotation data are updated weekly, guaranteeing that users get data most recently available. Data submitted by the user are stored in the database, where it can easily be updated, shared between users and exported in various formats. GeneTools provides three different tools: i) NMC Annotation Tool, which offers annotations from several databases like UniGene, Entrez Gene, SwissProt and GeneOntology, in both single- and batch search mode. ii) GO Annotator Tool, where users can add new gene ontology (GO) annotations to genes of interest. These user defined GO annotations can be used in further analysis or exported for public distribution. iii) eGOn, a tool for visualization and statistical hypothesis testing of GO category representation. As the first GO tool, eGOn supports hypothesis testing for three different situations (master-target situation, mutually exclusive target-target situation and intersecting target-target situation). An important additional function is an evidence-code filter that allows users, to select the GO annotations for the analysis. CONCLUSION: GeneTools is the first "all in one" annotation tool, providing users with a rapid extraction of highly relevant gene annotation data for e.g. thousands of genes or clones at once. It allows a user to define and archive new GO annotations and it supports hypothesis testing related to GO category representations. GeneTools is freely available through www.genetools.no

Algorithms↗

Comparison of maximum statistics for hypothesis testing when a nuisance parameter is present only under the alternative.

In many practical problems, a hypothesis testing involves a nuisance parameter which appears only under the alternative hypothesis. Davies (1977, Biometrika 64, 247-254) proposed the maximum of the score statistics over the whole range of the nuisance parameter as a test statistic for this type of hypothesis testing. Freidlin, Podgor, and Gastwirth (1999, Biometrics 55, 883-886) studied two other simpler maximum test statistics, the maximum of the score statistics at two extreme points of the nuisance parameter, and the maximum of the score statistics at three points of the nuisance parameter including the two extreme points. In this article, we compare the powers of these three maximum-type statistics in the context of three genetic problems.

Biometry↗

A hypothesis test for the end of a common source outbreak.

The objective of this article is to develop a hypothesis-testing procedure to determine whether a common source outbreak has ended. We consider the case when neither the calendar date of exposure to the pathogen nor the exact incubation period distribution is known. The hypothesis-testing procedure is based on the spacings between ordered calendar dates of disease onset of the cases. A simulation study was performed to evaluate the robustness of the methods to various models for the incubation period of infectious diseases. We investigated the impact of multiple testing on the overall outbreak-wise type I error probability. We derive expressions for the outbreak-wise type I error probability and show that multiple testing has minimal effect on inflating that error probability. The results are discussed in the context of the 2001 U.S. anthrax outbreak.

Anthrax↗

A new approach for interval estimation and hypothesis testing of a certain intraclass correlation coefficient: the generalized variable method.

We consider the problems of interval estimation and hypothesis testing for the intraclass correlation coefficient in an interrater reliability study when both raters and subjects are assumed to be randomly selected from populations of raters and subjects, respectively. We propose a novel approach for the confidence interval estimation and hypothesis testing using the concepts of generalized confidence interval (GCI) and generalized P-values. A simulation study is conducted to investigate the coverage probabilities of the GCI approach relative to the modified large sample (MLS) approach. Both methods tend to provide somewhat conservative coverage. Relative to the MLS approach, the GCI approach is closer to the correct (nominal) coverage for a two-sided interval, but farther to the correct coverage for a one-sided lower interval. Unlike the MLS approach, the GCI approach can also easily provide P-values. The fact that the GCI approach is suitable for confidence interval estimation and obtaining P-values makes the GCI approach a suitable candidate for making inference about interrater reliability.

Algorithms↗

Identity styles and hypothesis-testing strategies.

The author investigated the manner in which individuals with different identity styles evaluate information relevant to a trait hypothesis-testing task. Identity style refers to how individuals seek, process, evaluate, and use self-relevant information. In a sample of U.S. college students, the participants with normative identity styles displayed confirmatory biases that served to protect and preserve the hypothesis being evaluated. The students with informational identity styles also used confirmation-biased hypothesis-testing strategies. The participants with diffuse-avoidant styles did not discriminate between confirmatory and disconfirmatory information.

Adult↗

Using credibility intervals instead of hypothesis tests in SAGE analysis.

MOTIVATION: Statistical methods usually used to perform Serial Analysis of Gene Expression (SAGE) analysis are based on hypothesis testing. They answer the biologist's question: 'what are the genes with differential expression greater than r with P-value smaller than P?'. Another useful and not yet explored question is: 'what is the uncertainty in differential expression ratio of a gene?'. RESULTS: We have used Bayesian model for SAGE differential gene expression ratios as a more informative alternative to hypothesis tests since it provides credibility intervals. AVAILABILITY: The model is implemented in R statistical language script and is available under GNU/GLP copyleft at supplemental web site. SUPPLEMENTARY INFORMATION: http://www.ime.usp.br/~rvencio/SAGEci/

Algorithms↗

Experimental design and statistical methods for classical and bioequivalence hypothesis testing with an application to dairy nutrition studies.

Genetically modified (GM) corn hybrids have been recently compared against their isogenic reference counterparts in order to establish proof of safety as feedstuffs for dairy cattle. Most such studies have been based on the classical hypothesis test, whereby the null hypothesis is that of equivalence. Because the null hypothesis cannot be accepted, bioequivalence-testing procedures in which the alternative hypothesis is specified to be the equivalence hypothesis are proposed for these trials. Given a Type I error rate of 5%, this procedure is simply based on determining whether the 90% confidence interval on the GM vs. reference hybrid mean difference falls between two limits defining equivalence. Classical and bioequivalence power of test are determined for 4 x 4 Latin squares and double-reversal designs, the latter of which are ideally suited to bioequivalence studies. Although sufficient power likely exists for classical hypothesis testing in recent GM vs. reference hybrid studies, the same may not be true for bioequivalence testing depending on the equivalence limits chosen. The utility of observed or retrospective power to provide indirect evidence of bioequivalence is also criticized. Design and analysis issues pertain to Latin square and crossover studies in dairy nutrition studies are further reviewed. It is recommended that future studies should place greater emphasis on the use of confidence intervals relative to P-values to unify inference in both classical and bioequivalence-testing frameworks.

Analysis of Variance↗

Hypothesis testing with the similarity index.

Multilocus DNA fingerprinting methods have been used extensively to address genetic issues in wildlife populations. Hypotheses concerning population subdivision and differing levels of diversity can be addressed through the use of the similarity index (S), a band-sharing coefficient, and many researchers construct hypothesis tests with S based on the work of Lynch. It is shown in the present study, through mathematical analysis and through simulations, that estimates of the variance of a mean S based on Lynch's work are downwardly biased. An unbiased alternative is presented and mathematically justified. It is shown further, however, that even when the bias in Lynch's estimator is corrected, the estimator is highly imprecise compared with estimates based on an alternative approach such as 'parametric bootstrapping' of allele frequencies. Also discussed are permutation tests and their construction given the interdependence of Ss which share individuals. A simulation illustrates how some published misuses of these tests can lead to incorrect conclusions in hypothesis testing.

DNA Fingerprinting↗

Hypothesis-testing and nonlinguistic symbolic abilities in language-impaired children.

This study sought to clarify further the cognitive abilities of language-impaired children by examining their hypothesis-testing and nonlinguistic symbolic abilities. A discrimination learning task and a concept formation task were used to measure hypothesis-testing abilities, and a haptic (touch) recognition task was used to assess nonlinguistic symbolic abilities. Subjects were 10 language-impaired and 10 language-normal children matched for performance Mental Age. Measures of expressive and receptive language were also obtained from each child. The language-impaired children were found to perform significantly poorer than their MA controls on the haptic recognition task and on a portion of the discrimination learning task. No differences were found between the two groups' concept formation abilities. Correlational analyses revealed a particularly strong positive relationship between performance on the Peabody Picture Vocabulary Test and the haptic recognition tasks. It was speculated that this relationship was motivated by the symbolic demands of these tasks. One implication of this speculation is that a symbolic representational deficit might better explain the receptive language deficit than the expressive one.

Age Factors↗

Modulation of late alpha band oscillations by feedback in a hypothesis testing paradigm.

We used the electroencephalogram (EEG) to investigate whether positive and negative performance feedbacks differentially modulate late time-locked oscillatory brain activity in hypothesis testing. Ten college students serially tested hypotheses concerning a hidden rule by judging its presence or absence in triplets of digits, and revised them on the basis of an exogenous performance feedback. The EEG signal was convolved with a family of complex wavelets and induced brain potentials were extracted in the alpha range (8-13 Hz). The time-varying modulation of alpha activity time-locked to positive and negative feedback was analyzed in the 350-700 ms time-window. The results showed differential feedback-induced modulations of upper-alpha rhythms (> or =10 Hz) between 450 and 700 ms in parieto-occipital and central regions, and of lower-alpha rhythms (<10 Hz) between 350 and 450 ms in central regions. These results were interpreted in terms of differential functional roles of feedback in short-term memory and active inhibition/disinhibition of resources for subsequent hypothesis testing. Some implications for cognitive models of feedback are discussed.

Adult↗

Hypothesis testing of hazard ratio parameters in marginal models for multivariate failure time data.

Marginal hazard models for multivariate failure time data have been studied extensively in recent literature. However, standard hypothesis test statistics based on the likelihood method are not exactly appropriate for this kind of model. In this paper, extensions of the three commonly used likelihood hypothesis test statistics are discussed. Generalized Wald, generalized score and generalized likelihood ratio tests for hazard ratio parameters in a marginal hazard model for multivariate failure time data are proposed and their asymptotic distributions examined. The finite sample properties of these statistics are studied through simulations. The proposed method is applied to data from Busselton Population Health Surveys.

Age Factors↗

p values, hypothesis tests, and likelihood: implications for epidemiology of a neglected historical debate.

It is not generally appreciated that the p value, as conceived by R. A. Fisher, is not compatible with the Neyman-Pearson hypothesis test in which it has become embedded. The p value was meant to be a flexible inferential measure, whereas the hypothesis test was a rule for behavior, not inference. The combination of the two methods has led to a reinterpretation of the p value simultaneously as an "observed error rate" and as a measure of evidence. Both of these interpretations are problematic, and their combination has obscured the important differences between Neyman and Fisher on the nature of the scientific method and inhibited our understanding of the philosophic implications of the basic methods in use today. An analysis using another method promoted by Fisher, mathematical likelihood, shows that the p value substantially overstates the evidence against the null hypothesis. Likelihood makes clearer the distinction between error rates and inferential evidence and is a quantitative tool for expressing evidential strength that is more appropriate for the purposes of epidemiology than the p value.

Epidemiologic Methods↗