PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “construct validity”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Muscle testing response to provocative vertebral challenge and spinal manipulation: a randomized controlled trial of construct validity.

OBJECTIVE: To evaluate the relationship of muscle strength response to a provocative vertebral challenge and to spinal manipulation. DESIGN: Prospective double-blind randomized controlled trial: crossover and between subjects designs. SETTING: Laboratory: Center for Technique Research. PARTICIPANTS: Sixty-eight naive volunteers from the student body, staff and faculty of the college. INTERVENTIONS: Provocative vertebral challenge: standardized 4-5 kg force applied with a pressure algometer to the lateral aspects of the T3-12 spinous processes. INTERVENTION: manual high velocity low amplitude adjustment or switched-off activator sham. MAIN OUTCOME MEASURES: Piriformis muscle response was defined in two ways: reactivity (a decrease in muscle resistance, yes or nor, following a vertebral challenge); responsiveness (the cessation of reactivity following spinal manipulation). Relative response attributable to the maneuver (RRAM): the percent of an outcome attributable to the challenge or adjustment itself. RESULTS: Average RRAM = 16% reactivity to vertebral challenge; average RRAM = 0% responsiveness to spinal manipulation. Six to 10% of muscle tests were positive regardless of examiner, previous finding or intervention. CONCLUSIONS: For the population under investigation, muscle response appeared to be a random phenomenon unrelated to manipulable subluxation. In and of itself, muscle testing appears to be of questionable use for spinal screening and post-adjustive evaluation. Further research is indicated in more symptomatic populations, different regions of the spine, and using different indicator muscles.

Adult↗

Attitudes toward computers. A test of construct validity.

A study was undertaken to examine the factors that influence student and registered nurse attitudes toward computers. This article provides a report of the psychometric findings related to the use of Stronge and Brodt's Nurses' Attitudes Toward Computerization questionnaire. The convenience sample consisted of 394 registered nurses employed in a large metropolitan hospital and 299 baccalaureate student nurses. Principal component analysis was performed identifying three factors for the nurse sample, including nurses' work, organizational issues, and barriers to use. Four factors emerged in the student sample: nurses' work, barriers, organizational issues, and efficiency issues. These findings were inconsistent with previous studies and the underlying categories identified by Stronge and Brodt. Further development of the understanding of the substantive components of attitudes toward computers is recommended.

Adult↗

Evaluating the FONE FIM: Part I. Construct validity.

Rasch analysis was used in this paper to evaluate the Motor component of the FONE FIM, the telephone version of the Functional Independence Measure (FIM). For this purpose, 132 patients discharged from an inpatient geriatric assessment and rehabilitation program were assessed by trained research assistants using the FONE FIM. The results at 5 weeks post-discharge were compared to the observation FIMs (OBS FIMs) done at home 6 weeks post-discharge. These patients had an average age of 79 years and presented with multiple, complex medical problems and significant functional decline. The FONE FIM and the OBS FIM were shown to share a strikingly similar item hierarchy, based on Rasch item difficulty measures. Only bladder management and climbing stairs were misfitting items as indicated by item fit statistics. The same 13-item set and 4-point scales were shown to be psychometrically optimal for both the FONE FIM and the OBS FIM based on the person separation index. Further research is required to address the issue of the optimal item set and scale levels from psychometric and clinical perspectives.

Activities of Daily Living↗

Psychometric validation of a patient-reported sensory perception and preference instrument: the Sensory Perceptions Questionnaire.

This study evaluated the psychometric properties of content validity, construct validity, and test-retest reliability of a 23-item Sensory Perceptions Questionnaire (SPQ) used to survey sensory perceptions of intranasal corticosteroid sprays. Two patient cohorts (men and women aged > or =18 years who had at least a 1-year history of allergic rhinitis and had been using a corticosteroid nasal spray) were enrolled. The content validity and construct validity of the SPQ questions were evaluated using a cognitive debriefing method after cohort 1 (n=15) completed the SPQ. Test-retest reliability (assessed with intraclass correlation coefficients [ICCs] of the SPQ questions) was evaluated in cohort 2 (n=50), after they answered a Web-based version of the SPQ on two occasions, each separated by 7 days. In cohort 1, 7 of 15 patients believed all relevant sensory perceptions were addressed in the questionnaire. Although 8 patients mentioned at least 1 sensory perception that was not addressed, only 4 sensory perceptions were mentioned by more than 1 patient, and none was mentioned reliably by more than 2 patients. Those 4 sensory perceptions not addressed in the SPQ were all intentionally excluded, because they were potential symptoms of rhinitis or adverse events associated with intranasal corticosteroid spray use. Patients regarded the questions as straightforward, nonburdensome, and nonthreatening, signs suggesting the questions were not likely to challenge the construct validity of the SPQ. The responses to 2 questions (one in which patients were asked to indicate whether they were pleased or displeased overall with a particular spray; the other to indicate their overall product preference) were somewhat influenced by the effectiveness of the sprays. Results of test-retest reliability (cohort 2) showed both high (>0.8) and low (<0.7) ICCs. A high degree of correspondence between the 2 administrations produced a low between-patient variance, which likely resulted in lower ICCs. The SPQ adequately represents the sensory attributes reported by patients regarding intranasal corticosteroid spray use and, overall, is a valid measure of patient preference based on sensory perception.

Administration, Intranasal↗

Validation of the WHO Quality of Life assessment instrument (WHOQOL-100) in a population of Dutch adult psychiatric outpatients.

BACKGROUND: Research concerning the psychometric properties of the WHO Quality of Life Assessment Instrument (WHOQOL-100) in general populations of psychiatric outpatients has not been performed systematically. AIMS: To examine the content validity, construct validity, and reliability of the WHOQOL-100 in a general population of Dutch adult psychiatric outpatients. METHOD: A total of 533 psychiatric outpatients entered the study (438 randomly selected, 85 internally referred). Participants completed self-administered questionnaires for measuring quality of life (WHOQOL-100), psychopathological symptoms (SCL-90), and perceived social support (PSSS). In addition, they underwent two semi-structured interviews in order to obtain Axis-I and Axis-II diagnoses, according to DSM-IV. RESULTS: The drop-out percentage was low (7.1%). Of the 24 facets of the WHOQOL-100, 22 had a good distribution of scores, leaving out the facets physical environment and transport. Exploratory factor analysis revealed a four-factor structure, which was similar to earlier findings in patients with specific somatic diseases and depressive disorders. Various-a priori expected-positive and negative correlations were found between facets and domains of the WHOQOL-100, and dimensions of the SCL-90 and the PSSS-score, indicating good construct validity of the WHOQOL-100. The internal consistency of all facets and the four domains of the WHOQOL-100 was good (Cronbach's alpha's ranging from 0.62 to 0.93 and 0.64 to 0.84, respectively). Sparse and relatively low correlations were found between demographic characteristics (age and sex) and WHOQOL-100 scores. CONCLUSIONS: Content validity, construct validity, and reliability of the WHOQOL-100 in a population of adult Dutch psychiatric outpatients are good. The WHOQOL-100 appears to be a suitable instrument for measuring quality of life in adult psychiatric outpatients.

Adult↗

Short forms of the Child Perceptions Questionnaire for 11-14-year-old children (CPQ11-14): development and initial evaluation.

BACKGROUND: The Child Perceptions Questionnaire for children aged 11 to 14 years (CPQ11-14) is a 37-item measure of oral-health-related quality of life (OHRQoL) encompassing four domains: oral symptoms, functional limitations, emotional and social well-being. To facilitate its use in clinical settings and population-based health surveys, it was shortened to 16 and 8 items. Item impact and stepwise regression methods were used to produce each version. This paper describes the developmental process, compares the discriminative properties of the resulting four short-forms and evaluates their precision relative to the original CPQ11-14. METHODS: The item impact method used data from the CPQ11-14 item reduction study to select the questions with the highest impact scores in each domain. The regression method, where the dependent variable was the overall CPQ11-14 score and the independent variables its individual questions, was applied to the data collected in the validity study for the CPQ11-14. The measurement properties (i.e. criterion validity, construct validity, internal consistency reliability and test-retest reliability) of all 4 short-forms were evaluated using the data from the validity and reliability studies for the CPQ11-14. RESULTS: All short forms detected substantial variability in children's OHRQoL. The mean scores on the two 16-item questionnaires were almost identical, while on the two 8-item questionnaires they differed by only one score point. The mean scores standardized to 0-100 were higher on the short forms than the original CPQ11-14 (p < 0.001). There were strong significant correlations between all short-form scores and CPQ11-14 scores (0.87-0.98; p < 0.001). Hypotheses concerning construct validity were confirmed: the short-forms' scores were highest in the oro-facial, lower in the orthodontic and lowest in the paediatric dentistry group; all short-form questionnaires were positively correlated with the ratings of oral health and overall well-being, with the correlation coefficient being higher for the latter. The relative validity coefficients were 0.85 to 1.18. Cronbach's alpha and intraclass correlation coefficients ranged 0.71-0.83 and 0.71-0.77, respectively. CONCLUSION: All short forms demonstrated excellent criterion validity and good construct validity. The reliability coefficients exceeded standards for group-level comparisons. However, these are preliminary findings based on the convenience sampling and further testing in replicated studies involving clinical and general samples of children in various settings is necessary to establish measurement sensitivity and discriminative properties of these questionnaires.

Adolescent↗

[Development and validation of a modified quality of life questionnaire for patients treated by antireflux surgery].

INTRODUCTION: Past decade witnessed a growing interest in scientific medical publications on health related quality of life (HRQoL), which has yielded an increasing number of generic and disease-specific instruments. To date, several studies have evaluated the impact of GERD on HRQoL. AIMS: To develop a QoL questionnaire for patients with GERD underwent laparoscopic fundoplication (LF). This questionnaire was developed to be more comprehensive than existing measures. MATERIALS AND METHODS: We undertook a retrospective analysis of 116 patients underwent laparoscopic fundoplication for GERD between 1994 and 2002 in the 1st Department of Surgery, Semmelweis University. These patients--included 55 men and 61 women, with mean age of 46 years (14-77)--were used in the psychometric evaluation. Our questionnaire was developed using internationally accepted, valid QoL instruments and scales. RESULTS: Internal-consistency reliability was high (alpha value overall 0.95; dimensions 0.74-0.96). Using convergent and divergent validity, construct validity was evaluated by examining Pearson correlation coefficients between items and scales. Construct validity was demonstrated based on observed correlations. Known-groups validity of questionnaire is proved to. CONCLUSIONS: Our questionnaire is a short and user-friendly instrument. It has excellent reliability, construct and known-groups validity.

Adolescent↗

Psychometric documentation of a quality-of-life questionnaire for patients undergoing antireflux surgery (QOLARS).

BACKGROUND: The purpose of our study was to develop a quality-of-life (QoL) questionnaire for patients with gastroesophageal reflux disease (GERD) who have undergone laparoscopic fundoplication. This questionnaire was developed to be more comprehensive than existing measures. METHODS: Between 1994 and 2002, 252 patients underwent laparoscopic fundoplication for GERD in the 1st Department of Surgery, Semmelweis University. We undertook a retrospective analysis: each of 252 operated patients was given a questionnaire and was requested to complete it and return it in an enclosed envelope. A total of 116 patients returned completed questionnaires. The patients included 55 men and 61 women, with a mean age of 46 years (range 14-77). These patients were used in the psychometric evaluation. The questionnaire consisted of 50 questions (including the Visick score, EORTC-QLQ-C30, and a modified GERD-HRQL). RESULTS: Internal consistency reliability was high (alpha value overall, 0.95, range, 0.74-0.96). Using convergent and divergent validity, construct validity was evaluated by examining Pearson correlation coefficients between items and scales. Construct validity was demonstrated based on observed correlations. Known groups validity was upheld because patients who experienced more symptoms and patients who has higher Visick scores reported worse QoL than those with less symptoms or lower Visick scores. CONCLUSIONS: Our questionnaire is a short and user-friendly instrument with excellent psychometric properties. It has been found to be valid and reliable.

Adolescent↗

Validity of the MACTAR questionnaire as a functional index in a rheumatoid arthritis clinical trial. The McMaster Toronto Arthritis.

OBJECTIVES: The McMaster Toronto Arthritis patient preference questionnaire (MACTAR) is a functional index that measures change in impaired activities selected by each patient in a baseline interview, and change in rheumatoid arthritis (RA) disease activity. In addition, it contains questions on the state of physical, social and emotional function and overall health, and their relation to RA. We evaluated MACTAR's feasibility and validity (content validity, construct validity, and responsiveness). METHODS: A randomized trial of combined treatment in 155 patients with early RA; patients' mean age at baseline was 50 years and median disease duration since diagnosis was 4 months. RESULTS: Feasibility: MACTAR requires trained interviewers. In the trial, interviews took about 15 min. In longer term followup, activities selected at baseline may become less relevant as the pattern of disability changes. Followup from 153 patients (99%) was available. At least 5 impaired activities were identified and ranked by 147 patients (95%); interviewers could follow 99% of these. The scoring system proved complex and required amendments. Content validity: Although its main focus is physical function, the MACTAR also contains generic questions; 75% of the patients named at least one impaired activity from the category "mobility." Only 48% were covered by Health Assessment Questionnaire (HAQ) items. Construct validity: MACTAR scores correlate highly with other functional indices and with measures of disease activity. Responsiveness: At 16 weeks the standardized response mean for the total MACTAR score in the combined-treatment group was excellent, at 2.2. Items that directly address change were even more responsive. CONCLUSION: The MACTAR interview is a valid and highly responsive instrument to assess change in functional ability of patients with early RA with active disease. It provides insight into problems--mainly of physical function--that really matter to patients. For standard clinical trials and clinical care, feasibility of the MACTAR is limited and the simpler HAQ remains the instrument of choice.

Arthritis, Rheumatoid↗

Reliability, validity, and responsiveness of the American Shoulder and Elbow Surgeons subjective shoulder scale in patients with shoulder instability, rotator cuff disease, and glenohumeral arthritis.

BACKGROUND: Outcomes assessment after the treatment of shoulder disorders has involved the use of various condition-specific outcome instruments. The purpose of this study was to determine the psychometric properties of the American Shoulder and Elbow Surgeons subjective shoulder scale in patients with shoulder instability, rotator cuff disease, and glenohumeral arthritis. METHODS: Test-retest reliability, internal consistency, content validity, criterion validity, construct validity, and responsiveness to change were determined for the American Shoulder and Elbow Surgeons shoulder scale within subsets of an overall study population of 455 patients with shoulder instability, 474 patients with rotator cuff disease, and 137 patients with glenohumeral arthritis. RESULTS: There was acceptable test-retest reliability for the overall American Shoulder and Elbow Surgeons shoulder scale (intraclass correlation coefficient = 0.94) and ten of eleven domains. There was acceptable internal consistency for patients with instability (Cronbach alpha = 0.61), rotator cuff disease (0.64), and arthritis (0.62). There were acceptable floor and ceiling effects for patients with instability (0% and 1.3%, respectively), rotator cuff disease (0% for both), and arthritis (0% for both). There was acceptable and appropriate criterion validity, with significant correlations (p < 0.05) between the overall American Shoulder and Elbow Surgeons scale and the physical functioning, role-physical, and bodily pain domains of the Short Form-12 scale, and nonsignificant correlations (p > 0.05) with the role-emotional, mental health, vitality, and social function domains. There was acceptable construct validity, with all twenty-three hypotheses demonstrating significance (p < 0.05), and acceptable responsiveness to change for patients with instability (standardized response mean, 0.93), rotator cuff disease (1.16), and arthritis (1.11). CONCLUSIONS: The use of outcome instruments with psychometric properties that have been vigorously established is essential. The American Shoulder and Elbow Surgeons subjective shoulder scale demonstrated overall acceptable psychometric performance for outcomes assessment in patients with shoulder instability, rotator cuff disease, and glenohumeral arthritis.

Adolescent↗

Development and validation of a clinical scale for the diagnosis of drug-induced hepatitis.

The objective of this study is to present and validate a clinical scale for the diagnosis of drug-induced liver injury (DILI). Five components were selected to be included in the scale: temporal relationship between drug intake and the onset of clinical picture, exclusion of alternative causes, extrahepatic manifestations, rechallenge or accidental re-exposure, and previous report in medical literature. The relative importance of each component was weighed, and arbitrary scores were attributed. The probability of the diagnosis of DILI was expressed as a final score, which could vary from -6 to 20. Content validity, criterion validity, construct validity, and inter-rater reliability were studied. To analyze validity and reliability, a random sample of 50 cases of suspected DILI was drawn from a series of 120 cases reported to our unit. The classification of the 50 cases by three experts in DILI was used as the external standard in the study of criterion validity. Agreement between the scale and the standard, and agreement between two independent raters (inter-rater reliability) was analyzed by weighted kappa coefficient. There was agreement between the scale and the standard in 42 cases (84%) with a weighted kappa coefficient of 0.90. A good discriminatory capacity of the scale was found when construct validity was studied. Agreement between raters was observed in 86% of the cases, corresponding to the weighted kappa of 0.93. In conclusion, the clinical scale was shown to have a high-level of validity and inter-rater reliability as well as a good discriminatory capacity between different levels of probability. These data suggest that the scale is suitable for use in clinical practice and may contribute to overcome the difficulties in the process of causality assessment in DILI.

Bias↗

Reliability, validity, and responsiveness of the Lysholm knee scale for various chondral disorders of the knee.

BACKGROUND: The Lysholm knee scale is a condition-specific outcome measure that was originally designed to assess ligament injuries of the knee. The purpose of this study was to determine the psychometric properties of the Lysholm knee scale for various chondral disorders of the knee. METHODS: Test-retest reliability, internal consistency, content validity, criterion validity, construct validity, and responsiveness to change were determined for the Lysholm knee scale within subsets of an overall study population of 1657 patients with chondral disorders of the knee. The study population was a heterogeneous group of patients with various types of traumatic and degenerative chondral lesions, including isolated lesions and those associated with meniscal and ligament injuries. RESULTS: The overall Lysholm knee scale and six of the eight domains had acceptable test-retest reliability (intraclass correlation coefficient = 0.91) and internal consistency (Cronbach alpha = 0.65). The overall Lysholm knee scale demonstrated acceptable floor (0%) and ceiling (0.7%) effects; however, the floor effects for the domain of squatting and the ceiling effects for the domains of limp, instability, support, and locking were unacceptable (>30%). There was acceptable criterion validity with significant (p < 0.05) correlations between the overall Lysholm knee scale and the physical functioning, role-physical, and bodily pain domains of the Short Form-12 scale; the pain, stiffness, and function domains of the Western Ontario and McMaster Universities Osteoarthritis Index; and the Tegner activity scale. The overall Lysholm knee scale had acceptable construct validity, with all nine hypotheses demonstrating significance (p < 0.05), and it had acceptable responsiveness to change (effect size, 1.16; standardized response mean, 1.10), with large effects (> or = 0.80) for the domains of pain, limping, swelling, and squatting and a small effect (> or = 0.20) for the domain of instability. CONCLUSIONS: The Lysholm knee scale demonstrated overall acceptable psychometric performance for outcomes assessment of various chondral disorders of the knee, although some domains demonstrated suboptimal performance. Psychometric testing of other condition-specific knee instruments in patients with chondral disorders of the knee would be helpful to allow for comparison of psychometric properties.

Adult↗

Cardiovascular reactivity to psychological challenge: conceptual and measurement considerations.

OBJECTIVE AND METHODS: This article is a selective review of recent findings bearing on the conceptualization and measurement of cardiovascular reactivity to psychological challenge, with a focus on several issues relevant to the reliability, content validity, construct validity, and criterion validity of these measures. RESULTS AND CONCLUSIONS: With respect to reliability, use of standardized task demands and aggregated scores are associated with enhanced short-term reliability, but the long-term reliability of cardiovascular reactivity has not been sufficiently documented. With respect to content validity, existing evidence suggests that "vascular" or "cardiac" tasks may evoke responses that reflect similar distributions of individual difference, whereas associations between responses to "physical" and "psychological" tasks are modest. The evidence is not clear at present with respect to the importance of including affective or interpersonal stimuli as part of trait reactivity assessments. With respect to construct validity, existing data show that cardiovascular reactivity to psychological challenge is largely independent of standard measures of autonomic function. With respect to criterion validity, recent studies point to a number of methodological limitations that may have restricted our ability to detect lab-to-life generalizability of reactivity measures in the past. Continued progress in understanding and measuring reactivity as an individual difference dimension is essential in helping us to evaluate emerging evidence examining the relationship between reactivity and disease risk.

Cardiovascular Diseases↗

Reliability, validity, and responsiveness of the simple shoulder test: psychometric properties by age and injury type.

The purpose of this study was to measure the reliability, validity, and responsiveness of the Simple Shoulder Test (SST) scale and to examine these in patients stratified by age and injury type. Test-retest reliability, content validity, criterion validity, construct validity, and responsiveness were determined for the SST. The study population comprised 1077 patients with shoulder instability and rotator cuff injuries, ranging in age from 14 to 85 years. The SST demonstrated acceptable test-retest reliability (intraclass correlation coefficient >0.90) and content validity (floor and ceiling effects <10%). Correlations with the physical functioning component of the Short Form 12 were significant (r = 0.439, P < .05); however, the correlations were not significant when stratified by age group (>60 years) (r = 0.271, P = .349) and injury type (rotator cuff injury) (r = 0.337, P = .085). Correlations with the American Shoulder and Elbow Surgeons were also significant (r = 0.807, P < .001). The construct validity of the SST was acceptable, with all 8 hypotheses demonstrating significance (P < .05). The SST was responsive to change (effect size, 0.81; standardized response mean, 0.81). However, there were differences after stratification for age group and injury type. The SST demonstrated overall acceptable psychometric performance; however, differences were found when data were stratified by age and injury type.

Adolescent↗

The third factor of the WISC-III: it's (probably) not freedom from distractibility.

OBJECTIVE: This study examined the ecological validity, construct validity, and diagnostic utility of the third factor of the WISC-III, heuristically labeled "Freedom From Distractibility" (FFD). METHOD: A sample of 200 children, aged 6 to 11 years, with attention-deficit hyperactivity disorder (ADHD) completed the WISC-III, the Wide Range Achievement Test-Revised, and the Test of Variables of Attention. Objective parent and teacher report measures of attention and hyperactivity were completed. RESULTS: Mean FFD scores were significantly lower than other WISC-III factor scores. The diagnostic utility of FFD is limited, however, as the majority of these children did not show a significant relative weakness on this index. Correlational analyses failed to support the concurrent, ecological, or construct validity of the FFD. FFD scores were not correlated with a measure of sustained visual attention. Findings suggest that among children with ADHD, a low FFD score may be associated with the presence of a learning disability or poor academic performance. This finding was maintained after level of general intelligence was statistically controlled. CONCLUSIONS: Clinicians and researchers should not view FFD as a reliable or valid index of attention or as a diagnostic screening measure for identifying children with ADHD.

Achievement↗

Excessive reassurance seeking: delineating a risk factor involved in the development of depressive symptoms.

Six studies investigated (a) the construct validity of reassurance seeking and (b) reassurance seeking as a specific vulnerability factor for depressive symptoms. Studies 1 and 2 demonstrated that reassurance seeking is a reasonably cohesive, replicable, and valid construct, discernible from related interpersonal variables. Study 3 demonstrated that reassurance seeking displayed diagnostic specificity to depression, whereas other interpersonal variables did not, in a sample of clinically diagnosed participants. Study 4 prospectively assessed a group of initially symptom-free participants, and showed that those who developed future depressive symptoms (as compared with those who remained symptom-free) obtained elevated reassurance-seeking scores at baseline, when all participants were symptom-free, but did not obtain elevated scores on other interpersonal variables. Studies 5 and 6 indicate that reassurance seeking predicts future depressive reactions to stress. Taken together, the six studies support the construct validity of reassurance seeking, as well as its potential role as a specific vulnerability factor for depression.

Adolescent↗

Development and testing of a measure designed to assess the quality of care transitions.

BACKGROUND: To improve the quality of care delivered to older persons receiving care across multiple settings, interventions are needed. However, the absence of a patient-centred measure specifically designed to assess this care has constrained innovation. OBJECTIVE: To develop a rigorously designed and tested measure, the Care Transition Measure (CTM). SETTING: A large, integrated managed care organisation in Colorado with approximately 55,000 members over the age of 65 years. PARTICIPANTS: Patients 65 years and older who were recently discharged from hospital and received subsequent skilled nursing care in a facility or in the home. METHODS: Six focus groups of older persons and their caregivers (n=49) were established. Standard qualitative analytic techniques were applied to written transcripts and four key domains were identified: (1) information transfer; (2) patient and caregiver preparation; (3) self-management support; and (4) empowerment to assert preferences. Specific CTM items were developed, pilot tested, and refined. Psychometric testing, conducted in a different population but selected using the same entry criteria (n=60), included content and construct validity, intra-item variation, and floor/ceiling properties. RESULTS: Older patients and clinicians found the measure to be highly relevant and comprehensive (i.e. content validity). Construct validity was assessed by comparing items from the CTM to selected items from a measure developed by Hendriks and colleagues (Medical Care 2001; 39(3): 270-283). Inter-item Spearman correlations ranged 0.388-0.594. No significant floor or ceiling effects were detected. CONCLUSIONS: The CTM was developed with substantial input from older patients and their caregivers. Psychometric testing suggested that the measure was valid. The CTM may serve to fill an important gap in health system performance evaluation by measuring the quality of care delivered across settings.

Journal Article↗

The parent-form Child Health Questionnaire in Australia: comparison of reliability, validity, structure, and norms.

OBJECTIVE: To improve the ability to describe and compare child health within and between countries, using standardized multidimensional child health measures. METHODS: Data on population-specific psychometrics, the measurement structure, and norms are a vital prerequisite. These properties for the Child Health Questionnaire (CHQ) were examined for an Australian population and compared with the originating U.S. data. The CHQ 50-item parent-report was completed by 5,414 parents of children aged 5-18 years. Multi-item/multi-trait analysis tested convergent and discriminatory validity. Construct validity, test-retest reliability, comparative population mean scale scores, and the summary score factor structure were examined. RESULTS: Item and scale internal consistency and item discriminant validity results were good to excellent, and construct (concurrent) validity was supported. Australian children had higher scores than U.S. children except for Family Activities and Physical Functioning. The factor structure of the two summary scores for American children was not replicated in the normative sample but held for a subsample of children with one or more health conditions. CONCLUSIONS: The CHQ PF50 performed well in Australia at item and scale level. However, the physical and psychosocial summary scores are not supported for population-level analyses but may be of value for sub-groups of children with health problems.

Adolescent↗