PubMed Health⌕ Search

Biomedical subjects

K F Cook

Publications and source records attributed to K F Cook.

7 recordsLinked to original sources

A psychometric analysis of the measurement level of the rating scale, time trade-off, and standard gamble.

A fundamental assumption of utility-based analyses is that patient utilities for health states can be measured on an equal-interval scale. This assumption, however, has not been widely examined. The objective of this study was to assess whether the rating scale (RS), standard gamble (SG), and time trade-off (TTO) utility elicitation methods function as equal-interval level scales. We wrote descriptions of eight prostate-cancer-related health states. In interviews with patients who had newly diagnosed, advanced prostate cancer, utilities for the health states were elicited using the RS, SG, and TTO methods. At the time of the study, 77 initial and 73 follow-up interviews had. been conducted with a consecutive sample of 77 participants. Using a Rasch model, the boundaries (Thurstone Thresholds) between four equal score sub-ranges of the raw utilities were mapped onto an equal-interval logit scale. The distance between adjacent thresholds in logit units was calculated to determine whether the raw utilities were equal-interval. None of the utility scales functioned as interval-level scales in our sample. Therefore, since interval-level estimates are assumed in utility-based analyses, doubt is raised regarding the validity of findings from previous analyses based on these scales. Our findings need to be replicated in other contexts, and the practical impact of non-interval measurement on utility-based analyses should be explored. If cost-effectiveness analyses are not found to be robust to violations of the assumption that utilities are interval, serious doubt will be cast upon findings from utility-based analyses and upon the wisdom of expending millions in research dollars on utility-based studies.

Cost-Benefit Analysis↗

SODA (severity of dyspepsia assessment): a new effective outcome measure for dyspepsia-related health.

The aim of this research was to develop and evaluate an instrument for measuring dyspepsia-related health to serve as the primary outcome measure for randomized clinical trials. Building on our previous work we developed SODA (Severity of Dyspepsia Assessment), a multidimensional dyspepsia measure. We evaluated SODA by administering it at enrollment and seven follow-up visits to 98 patients with dyspepsia who were randomized to a 6-week course of omeprazole versus placebo and followed over 1 year. The mean age was 53 years, and six patients (6%) were women. Median Cronbach's alpha reliability estimates over the eight visits for the SODA Pain Intensity, Non-Pain Symptoms, and Satisfaction scales were 0.97, 0.90, and 0.92, respectively. The mean change scores for all three scales discriminated between patients who reported they were improved versus those who were unchanged, providing evidence of validity. The effect sizes for the Pain Intensity (.98) and Satisfaction (.87) scales were large, providing evidence for responsiveness. The effect size for the Non-Pain Symptoms scale was small (.24), indicating lower responsiveness in this study sample. SODA is a new, effective instrument for measuring dyspepsia-related health. SODA is multidimensional and responsive to clinically meaningful change with demonstrated reliability and validity.

Adult↗

The measurement level and trait-specific reliability of 4 scales of shoulder functioning: an empiric investigation.

OBJECTIVE: To evaluate 4 scales of shoulder function with respect to (1) their precision at different levels of shoulder function and (2) the measurement level of their raw scores (interval vs ordinal). DESIGN: Partial credit model calibration. SETTING: Office of private practice orthopedic surgeon with practice limited to the shoulder. PARTICIPANTS: One-hundred ninety-two shoulder patients. INTERVENTIONS: Participants completed the American Shoulder and Elbow Surgeons Patient Self-Evaluation Form (function subscale, modified), the disability subscale of the Shoulder Pain and Disability Index, the Simple Shoulder Test, and the function subscale of the University of Pennsylvania Shoulder Scale. MAIN OUTCOME MEASURES: The patients' responses were calibrated by using a partial credit model. We calculated standard errors of measurement and plotted the 95% confidence interval for different levels of shoulder functioning. We compared scales' raw scores with their equal interval measures obtained in the Rasch calibration. RESULTS: The scales did not measure all levels of shoulder functioning with equal precision, suggesting that commonly used reliability estimates misrepresent scale precision in certain subpopulations. CONCLUSIONS: The scales' raw scores were found to be not of equal interval, calling into question the scoring systems recommended by the developers of these scales and the use of the scores in some statistical procedures.

Adult↗

Comparison of the University of California-Los Angeles Shoulder Scale and the Simple Shoulder Test with the shoulder pain and disability index: single-administration reliability and validity.

BACKGROUND AND PURPOSE: Shoulder scales are often used to evaluate treatment efficacy, yet little is known about the psychometric properties of these scales. Only one scale has undergone psychometric scrutiny: the Shoulder Pain and Disability Index (SPADI). This study compared 2 shoulder measures-the University of California-Los Angeles (UCLA) Shoulder Scale and the Simple Shoulder Test (SST)-with the SPADI. SUBJECTS: One hundred ninety-two patients with shoulder disorders were recruited from one physician's office to complete the self-report sections of the 3 scales. METHODS: Cronbach alpha values and standard errors of measurement (SEM) were calculated for each of the multi-item subscales. Validity was examined through calculation of correlation coefficients among the 3 scales. Factor analysis was completed to assess the underlying constructs of the SPADI and the SST. RESULTS: Cronbach alpha values ranged from.85 to.95. The SEM values for the multi-item scales ranged from 4.75 to 11.65. Evidence for validity to reflect function was indicated by the correlation between the SST and the SPADI disability subscale. The factor analysis of the SPADI revealed loading on 1 factor, whereas the SST loaded on 2 factors. CONCLUSION AND DISCUSSION: All scales demonstrated good internal consistency, suggesting that all items for each scale measure the same construct. However, the SEMs for all scales were high. Factor loading was inconsistent, suggesting that patients may not distinguish between pain and function.

Adolescent↗

Evaluation of a multidimensional measure of dyspepsia-related health for use in a randomized clinical trial.

In previous work, we developed a multidimensional measure of dyspepsia-related health. To evaluate the adequacy of this instrument as an outcome measure for a large-scale, multicenter, randomized clinical trial, we used Rasch analysis to address three questions: (1) Are the scales interval-level? (2) Do the scales measure precisely across the entire range of dyspepsia outcomes? (3) Do the scales' items have an optimal number of response categories? We found that the scales were not interval-level and that they did not measure effectively at low or high levels of the dyspepsia-related outcomes. Our results also suggest that patients were capable of discriminating among only four- to seven-item response categories. Further studies are needed to identify items that effectively measure high and low levels of dyspepsia-related outcomes and to validate that decreasing the number of response categories improves the psychometric properties of these scales.

Adult↗

A comparison of three polytomous item response theory models in the context of testlet scoring.

An alternative to dichotomous scoring of multiple items anchored to a common stem is scoring these items as a single polytomous item (testlet scoring). This study systematically compared the partial credit model (PCM), the generalized partial credit model (GPCM), and the graded response model (GRM) in the context of testlet scoring. Data sets included a sample from the fall 1994 administration of the SAT I (N = 2,548) and a simulated data set. Theta estimation, information, and model fit were analyzed. Correlations among theta estimates ranged from 0.9748 to 0.9921. The relationship among the information functions of the PCM, GPCM and the GRM reflected the discrimination parameter estimates for the latter two models. Suggestions are made with regard to model selection.

Aptitude Tests↗

Parameter recovery for the partial credit model using MULTILOG.

This study investigated parameter recovery for the partial credit model using the MULTILOG computer program. Factors studied were the sample size and the number of item parameters, which were manipulated by systematically varying the number of steps per item and the number of items. The findings suggest that the ratio of sample size to number of item parameters being estimated as a "rule of thumb" can be a more complete guideline when the number of steps per item is taken into account. Accurate estimation of ability can be obtained across all conditions, even with sample sizes as small as 250. With regard to estimation of step values, however, more caution is warranted. Accurate estimation of the step values of items which have more categories requires larger sample sizes for a given number of total parameters to be estimated.

Humans↗