PubMed Health⌕ Search

Biomedical subjects

Richard A Charter

Publications and source records attributed to Richard A Charter.

16 recordsLinked to original sources

Psychometric properties of the Folstein Mini-Mental State Examination.

Criterion-referenced (Livingston) and norm-referenced (Gilmer-Feldt) techniques were used to measure the internal consistency reliability of Folstein's Mini-Mental State Examination (MMSE) on a large sample (N = 418) of elderly medical patients. Two administration and scoring variants of the MMSE Attention and Calculation section (Serial 7s only and WORLD only) were investigated. Livingston reliability coefficients (rs) were calculated for a wide range of cutoff scores. As necessary for the calculation of the Gilmer-Feldt r, a factor analysis showed that the MMSE measures three cognitive domains. Livingston's r for the most widely used MMSE cutoff score of 24 was .803 for Serial 7s and .795 for WORLD. The Gilmer-Feldt internal consistency reliability coefficient was .764 for Serial 7s and .747 for WORLD. Item analysis showed that nearly all of the MMSE items were good discriminators, but 12 were too easy. True score confidence intervals should be applied when interpreting MMSE test scores.

Aged↗

Psychometric properties of the benton visual form discrimination test.

Coefficient alpha and an item analysis were calculated for the 16-item Benton Visual Form Discrimination Test (VFDT) using a heterogeneous sample (N = 293) of mostly elderly medical patients who were suspected of having cognitive impairment. The total score reliability was .74. An item analysis found that 15 of the items were within established criteria for item difficulty, however, 5 items were found to be poor discriminators. Through the use of confidence intervals around observed scores, it was shown that the current classification criterion for the VFDT demands a higher reliability coefficient than what was found. Also, evidence for the test's insufficient level of difficulty is presented. It is difficult to recommend this test for clinical use.

Adult↗

Evidence for aging as the cause of Alzheimer's disease.

Part 1 presents the results of a meta-analytic study on the effects of aging on intelligence. Analysis of a total of 20 longitudinal samples shows that most of the intelligence scores rose before the age of 50 and fell at a progressively increasing rate after the age of 50. An equation describing this rise and fall in intelligence was derived. Part 2 shows the relationship between the predicted prevalence of Alzheimer's disease (from the equation derived in Part 1) and the prevalence of the disease obtained from 10 studies. The predictive curve fit so well with the observed prevalence data that the results can be interpreted as evidence that Alzheimer's is a manifestation of aging.

Aged↗

A cautionary note on deviation scores for the WISC-III and WAIS-III.

This article highlights some dangers inherent in interpreting individual examinee strengths and weaknesses in WAIS-III and WISC-III profiles. In both manuals, there are tables providing point estimates for determining if a significant difference exists when comparing one subtest to the average of several subtests. However, these point estimates may lead to interpretation errors. A confidence interval approach that provides a solution to the interpretation problem is described and illustrated. These intervals show that differences between a single subtest score and the average subtest score that just reach statistical significance carry the potential of having no clinical significance.

Cognition↗

MMPI-2: Confidence intervals for random responding to the F, F Back, and VRIN scales.

Several studies have investigated random responding to the F, F Back, and VRIN scales. Only one study attempted to provide practical cutoff scores for these scales, but was unable to reach definitive cutoffs. This study uses the normal approximation to the binomial distribution and provides confidence interval bounds for random responding at the 95, 90, and 85% levels for the F, F Back, and VRIN scales. The possibility that humans asked to respond randomly produce F, F Back, and VRIN scores different from computer-generated random scores was investigated. The results show nonsignificant differences between human and computer responses for the F and F Back scales and mixed results for the VRIN scale.

Computer Simulation↗

Estimating the reliability of a test split into two parts of equal or unequal length.

When the reliability of test scores must be estimated by an internal consistency method, partition of the test into just 2 parts may be the only way to maintain content equivalence of the parts. If the parts are classically parallel, the Spearman-Brown formula may be validly used to estimate the reliability of total scores. If the parts differ in their standard deviations but are tau equivalent, Cronbach's alpha is appropriate. However, if the 2 parts are congeneric, that is, they are unequal in functional length or they comprise heterogeneous item types, a less well-known estimate, the Angoff-Feldt coefficient, is appropriate. Guidelines in terms of the ratio of standard deviations are proposed for choosing among Spearman-Brown, alpha, and Angoff-Feldt coefficients.

Analysis of Variance↗

A breakdown of reliability coefficients by test type and reliability method, and the clinical implications of low reliability.

The author presented descriptive statistics for 937 reliability coefficients for various reliability methods (e.g., alpha) and test types (e.g., intelligence). He compared the average reliability coefficients with the reliability standards that are suggested by experts and found that most average reliabilities were less than ideal. Correlations showed that over the past several decades there has been neither a rise nor a decline in the value of internal consistency, retest, or interjudge reliability coefficients. Of the internal consistency approaches, there has been an increase in the use of coefficient alpha, whereas use of the split-half method has decreased over time. Decision analysis and true-score confidence intervals showed how low reliability can result in clinical decision errors.

Decision Support Techniques↗

Study samples are too small to produce sufficiently precise reliability coefficients.

In a survey of journal articles, test manuals, and test critique books, the author found that a mean sample size (N) of 260 participants had been used for reliability studies on 742 tests. The distribution was skewed because the median sample size for the total sample was only 90. The median sample sizes for the internal consistency, retest, and interjudge reliabilities were 182, 64, and 36, respectively. The author presented sample size statistics for the various internal consistency methods and types of tests. In general, the author found that the sample sizes that were used in the internal consistency studies were too small to produce sufficiently precise reliability coefficients, which in turn could cause imprecise estimates of examinee true-score confidence intervals. The results also suggest that larger sample sizes have been used in the last decade compared with those that were used in earlier decades.

Humans↗

Boston Naming Test: problems with administration and scoring.

The poorly written administration and scoring instructions for the Boston Naming Test allow too wide a range of interpretations. Three different, seemingly correct interpretations of the scoring methods were compared. The results show that these methods can produce large differences in the total score.

Female↗

Combining reliability coefficients: possible application to meta-analysis and reliability generalization.

Formulae for combining reliability coefficients from any number of samples are provided. These formulae produce the exact reliability one would compute if one had the raw data from the samples. Needed are the sample means, standard deviations, sample sizes, and reliability coefficients. The formulae work for coefficient alpha, KR-20, retest, alternate-forms, split-half, interrater (intraclass), Gilmer-Feldt, Angoff-Feldt, validity, and other coefficients. They may be particularly useful for meta-analytic and reliability generalization studies.

Humans↗

Note on Knight's analysis of the WAIS-III instruction effect on the Matrix Reasoning subtest.

Knight's 2003 analysis of the effect of the WAIS-III instructions on the Matrix Reasoning subtest was based on multiple t tests, which is a violation of conventional statistical procedures. Using this procedure significant differences were found between the group who know the subtest was untimed versus the group which did not know if the subtest was timed or untimed. Reanalysis of the data used three statistical alternatives: (a) Bonferroni correction for all possible t tests, (b) one-way analysis of variance, and (c) selected t tests with the Bonferroni correction. All three analyses yielded nonsignificant differences between means, thereby changing the conclusions of Knight's study.

Decision Making↗

Millon Clinical Multiaxial Inventory (MCMI-III): the inability of the validity conditions to detect random responders.

The effectiveness of the MCMI-III Validity scale, Scale X, and the Clinical Personality Pattern scales to detect random responding is put to the test. The binomial expansion and Monte Carlo techniques were used. If the examiner is willing to interpret tests of questionable validity, then 50% of the random responders will not be detected. Scale X and the Clinical Personality Pattern scales were useless in detecting random responders.

Adult↗