PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Validity”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Impact of the validator and the validation method on the outcome of occlusal caries diagnosis.

In studies evaluating the performance of caries-diagnostic methods, a validation of the true state of disease is needed. The aim of this study was to evaluate the impact of the validator and the validation system (stereomicroscopy or radiography of tooth sections) on the outcome of diagnostic tests for occlusal caries. The material consisted of 60 extracted third molars which were serially sectioned (500-600 microns thick). Four observers examined the sections by two caries validation methods: stereomicroscopy and film radiography. The presence of caries in the occlusal surfaces of these teeth had previously been recorded by visual inspection and conventional film radiography. The kappa values for interobserver agreement within one validation method ranged from 0.44 to 0.76 for radiography and from 0.47 to 0.60 for microscopy. The intraobserver agreement with the two methods was low (range 0.31-0.49), and by cross-tabulating the data, it was found that the disagreements originated in a consistently deeper lesion score with stereomicroscopy than with radiography by all observers except 1. By use of receiver operating characteristic curve areas, little impact of the validation method was seen when visual inspection was validated (against microscopy mean area = 0.75, against radiography mean area = 0.74). The mean receiver operating characteristic area was higher when diagnostic film radiographs were validated against radiography (0.68) than against microscopy (0.63). The differences between observers within a validation method were larger with microscopy than with radiography. In conclusion, caries validation methods are subject to variability. The outcome of caries--diagnostic tests may be influenced both by the validator and the validation method.

Dental Caries↗

The measurement of instrumental ADL: content validity and construct validity.

A new measure of Instrumental Activities of Daily Living (IADL), which is able to discriminate among the large group of elderly who do not depend on help, was tested for content validity and construct validity. Most assessments of functional ability include Physical ADL (PADL) and Instrumental ADL (IADL). PADL-scales assess the basic capacity of persons to care for themselves. IADL-scales are used to assess somewhat higher levels of performance, such as the ability to perform household chores or go shopping. Data were collected from 734 70-year-old people in Denmark in the county of Copenhagen. The measure of Instrumental ADL included 30 activities in relation to tiredness and reduced speed. Construct validity was tested by the Rasch model for item analysis; internal validity was specifically addressed by assessing the homogeneity of items under different conditions. The Rasch item analysis of IADL showed that 14 items could be combined into two qualitatively different additive scales. The IADL-measure complies with demands for content validity, distinguishes between what the elderly actually do, and what they are capable of doing, and is a good discriminator among the group of elderly persons who do not depend on help. It is also possible to add the items in a valid way. However, to obtain valid IADL-scales, we omitted items that were highly relevant to especially elderly women, such as house-work items. We conclude that the criteria employed for this IADL-measure are somewhat contradictory.

Activities of Daily Living↗

Validation of dietary intakes of protein and energy against 24 hour urinary N and DLW energy expenditure in middle-aged women, retired men and post-obese subjects: comparisons with validation against presumed energy requirements.

OBJECTIVES: To compare validation of reported dietary intakes from weighed records against urinary nitrogen excretion and energy expenditure measured by DLW, and to examine the utility of the Goldberg cut-off for EI:BMR in the identification of under-reporters. DESIGN: Energy (EI) and nitrogen (protein) intake (NI) were measured by 16 d of weighed diet records collected over 1 y. They were validated against urinary nitrogen excretion in 5-8 (mean 6.0) 24 h urine collections and total energy expenditure (EE) measured by doubly labelled water (DLW). Basal metabolic rate (BMR) as measured by whole body calorimetry in women or bedside ventilated hood (Deltatrac) in men. Individual subjects were identified as under-reporters if Urine N:NI was > 1.00 or if EI:EE was < 0.79. The agreement between the two ratios in detecting under-reporting was examined. The results from the direct validation by DLW were also compared with validation using the Goldberg cut-off for EI:BMR (Goldberg et al, 1991). SUBJECTS: Eighteen women aged 50-65 y and 27 men aged 55-87 y were selected from participants in two larger dietary surveys as representing the full range of dietary reporting as measured by Urine N:NI. Data from a previous study of 11 post-obese subjects were also included. RESULTS: The two ratios, Urine N:NI and EI:EE, were significantly related (r = -0.48, P < 0.01). Using the above cut-offs, seven (4F, 3M) subjects were identified as under-reporters by both methods, one (1M) by Urine N:NI only and 8 (3F, 5M) by EI:EE only. There was close agreement in post-obese subjects where 6 subjects showed a substantial degree of under-reporting by both methods (r = -0.87, P < 0.001). The correlation between direct validation by DLW and EI:BMRest was 0.65 (P < 0.001). Some limitations of the Goldberg cut-off for identifying individual under-reporters were demonstrated. CONCLUSIONS: EI:EE provides an estimate of the degree of under-reporting of energy at the group and individual level. Urine N:NI identifies under-reporting of protein intake and the most obvious under-reporters of energy, but is probably of lesser value in estimating the overall degree of under-reporting of energy at group level. Good validation by EI:BMR depends on knowledge of physical activity at both group and individual level. However, the correlation of 0.65 between EI:EE and EI:BMRest suggests that EI:BMR could be usefully incorporated into analysis of data from epidemiological studies. Validation measures consisting of at least predicted EI:BMR ratios and urinary measures should be incorporated into dietary surveys. SPONSORSHIP: This work was funded by the Ministry of Agriculture Fisheries and Food, the Medical Research Council, the Cancer Research Council and the Swedish Medical Research Council and the Henning and Johan Throne-Holst Foundation.

Aged↗

Validation and cross-validation of the PTSD subscale of the MMPI with civilian trauma victims.

The 49-item MMPI PTSD Subscale, developed and validated with Vietnam combat veterans, was administered to validation and cross-validation samples of Posttraumatic Stress Disorder (PTSD) patients who had experienced non-military traumatic events and to psychiatric controls (total N = 69). Using a cutting score of 19, derived from the validation sample only, the PTSD subscale correctly classified 87% of all validation subjects and 88% of all cross-validation subjects. Results strongly support the utility of MMPI assessment of PTSD with civilian trauma victims as one component of a broad assessment strategy.

Accidents↗

The validation of three human reliability quantification techniques--THERP, HEART and JHEDI: Part II--Results of validation exercise.

This is the second of three papers dealing with the validation of three Human Reliability Assessment (HRA) techniques. The first paper introduced the need for validation, the techniques themselves and pertinent validation issues. This second paper details the results of the validation study carried out on the Human Reliability Quantification techniques THERP, HEART and JHEDI. The validation study used 30 real Human Error Probabilities (HEPs) and 30 active Human Reliability Assessment (HRA) assessors, 10 per technique. The results were that 23 of the assessors showed a significant correlation between their estimates and the real HEPs, supporting the predictive accuracy of the techniques. Overall precision showed 72% (60-87%) of all HEPs to be within a factor of 10 of the true HEPs, with 38% of all estimates being within a factor of three of the true values. Techniques also tended to be pessimistic rather than optimistic, when they were imprecise. These results lend support to the empirical validity of these three approaches.

Benchmarking↗

Validation of the NOSGER (Nurses' Observation Scale for Geriatric Patients): reliability and validity of a caregiver rating instrument.

The Nurses' Observation Scale for Geriatric Patients (NOSGER) is a rating scale for use in geriatric patients that can be applied by nurses or other caregivers. It deals with the daily behavior of elderly patients and measures impairment in six areas (dimensions): memory; instrumental activities of daily living (IADL); (basic) activities of daily living (ADL); mood; social behavior; and disturbing behavior. Objectivity, stability, construct validity, and acceptance of the scale have been established in previous studies using an earlier version of the NOSGER. The present validation study considered 50 healthy old subjects, 25 patients with mild dementia, 25 patients with advanced (mostly moderate according to DSM-III-R criteria) dementia, and 25 elderly patients with depression. The NOSGER was completed by relatives in the case of subjects living in their own homes and by nurses or other caregivers for institutionalized subjects. In addition to the NOSGER, selected tests of concentration, memory, and performance were applied as outside criteria. Interrater reliability (objectivity) was estimated by variance component analysis. Values between rtt = .68 and rtt = .89 (all p < .001) were found for the six NOSGER dimensions, the values being higher for the cognitive dimensions (memory, IADL, ADL) than for the noncognitive ones (mood, social behavior, disturbing behavior). Retest reliability (stability), which was calculated via rank order correlations, was somewhat higher for the cognitive NOSGER dimensions (memory rs = .91, IADL rs = .92, ADL rs = .88; p < .001) than for the noncognitive ones (mood rs = .85, social behavior rs = .87, disturbing behavior rs = .84; p < .001). All these values satisfy the level of rtt > or = .80 required in accordance with psychometric standards. The concurrent validity of the NOSGER dimensions was assessed using correlations with external criteria with which similarity of content was expected. The NOSGER dimensions memory, IADL, ADL, and social behavior were found to correlate closely with external criteria of similar content, whereas no satisfactory concurrent validities were found for the dimensions mood or disturbing behavior. The NOSGER dimensions were also correlated with a number of unrelated external criteria so as to reveal any discordances. For the dimensions memory, IADL, ADL, and social behavior, no clear-cut discriminant validities were found. This suggests that these four dimensions may function as parameters not just of different areas of behavior, but also of a general factor that might be described as "cognitive intactness." As a further aspect of construct validity, significant differences (all p < .001) between the four groups of subjects were found in five of the six NOSGER dimensions (memory, IADL, ADL, mood, social behavior): The healthy subjects differed significantly from all three patient groups in five of the six dimensions; the moderately demented group differed from the depressed group in four of the six dimensions and from the mildly demented group in two of the six dimensions; and the mildly demented group differed significantly from the depressed group in terms of mood (significance levels are after application of the Bonferroni correction). Significant group differences (p generally < .001) were also found for most of the objective performance tests used (data not presented).

Activities of Daily Living↗

Internal validity of attention deficit hyperactivity disorder, oppositional defiant disorder, and overt conduct disorder symptoms in young children: implications from teacher ratings for a dimensional approach to symptom validity.

Uses a dimensional approach to evaluate the internal validity of the attention deficit hyperactivity disorder (ADHD) inattention (I) and hyperactivity/impulsivity (H/I), oppositional defiant disorder (ODD), and overt conduct disorder (CD) symptoms (i.e., whether a symptom has a stronger correlation with its own dimension than the other three dimensions). In Study 1, teachers rated 1,445 children on the DSM-III-R I, H/I, ODD, and overt CD symptoms. In Study 2, teachers rated 1,711 children on the DSM-IV I, H/I, ODD, and overt CD symptoms. All the I symptoms showed internal validity in both studies. In contrast, the H/I symptoms and the ODD symptoms, especially the H/I symptoms, showed weaker internal validity. All the overt CD symptoms showed internal validity except the DSM-IV bullies others symptom, with this symptom being more strongly related to the ODD dimension. Confirmatory factor analysis provided support for a 4-factor model consisting of I, H/I, ODD, and overt CD factors. Finally, the importance of internal validity for the construct validation of the disruptive behavior symptoms is discussed.

Adolescent↗

[Translation and validation of the Revised Social Anhedonia Scale (SAS Social Anhedonia Scale, M.L. Eckblad, L.J. Chapman et al., 1982). Study of the internal and concurrent validity in 126 normal subjects].

The Revised Social Anhedonia Scale (SAS) with 40 items (Eckblad et al., 1982) which studies the social dimension of anhedonia has been validated in the United-States (Mishlove & Chapman, 1985). However, no french translation and validation of this scale has been made to date. This work presents the french translation of the Social Anhedonia Scale and its validation. After a back-translation and final adjustment, it has been submitted to a sample of 126 control subjects from the general population. Furthermore, they were asked to fill two other scales: the Chapman Physical Anhedonia (PAS) with 61 items and the Fawcett Pleasure Scale (36 items), both of them exploring the subjects answer in terms of anhedonia/hedonia towards social, sensorial and/or physical experiences. The internal validity has been determined on the one hand by the Cronbach alpha coefficient which showed a strong unidimensional characteristic (0.80) and the other hand by the correlation of each item with the total score using the point biserial coefficient which ranged from .204 to .559. The concurrent validation has been determined by the Pearson correlation coefficient between the french version of social anhedonia scale and the french version of physical anhedonia scale. The values were .42, p = .001. Furthermore, this two scales are significant inversely correled to the french version of the pleasure scale: r = -.22, p = .0125 for the first, and r = -.26, p = .0027 for the second. The internal and concurrent validity of the french version of the revised social anhedonia scale should allow to improve our understanding of anhedonia in psychiatry and psychopathology.

Adolescent↗

[Internal consistency, factorial validity and discriminant validity of the French version of the psychological demands and decision latitude scales of the Karasek "Job Content Questionnaire"].

BACKGROUND: The job demands-control model developed by Karasek has greatly influenced research on psychosocial factors at work and health. Validity of the English version of the psychological demands and decision latitude scales is documented. Psychometric qualities of the French version are investigated here in a representative sample of the general population, including blue-collars and white-collars. METHODS: The French translation of the psychological demand and decision latitude scales was administered by interview in a representative sample of the Quebec working population (N = 1,110). Internal consistency and factorial validity of the instrument were studied among white-collars and blue-collars separately. Discriminant validity was assessed for the whole population. RESULTS: Cronbach alpha coefficients, varying between 0.68 and 0.85, support the internal consistency of the scales. Demographic distribution of the scales and intercorrelations were consistent with the English version. Results of the factor analysis were consistent with the two dimensions expected from the theory. Mean scale scores and variations in the prevalence of high psychological demands combined with low decision latitude by age, sex, education, and job category support the discriminant validity of the instrument. CONCLUSIONS: Results support internal consistency, factorial validity, and discriminant validity of the French version of the psychological demands and decision latitude scales of the Karasek "Job Content Questionnaire" for white-collars and for blue-collars of the general population.

Adolescent↗

Victoria Symptom Validity Test: efficiency for detecting feigned memory impairment and relationship to neuropsychological tests and MMPI-2 validity scales.

Error scores and response times from a computer-administered, forced-choice recognition test of symptom validity were evaluated for efficiency in detecting feigned memory deficits. Participants included controls (n = 95), experimental malingerers (n = 43), compensation-seeking patients (n = 206), and patients not seeking financial compensation (n = 32). Adopting a three-level cut-score system that classified participant performance as malingered, questionable, or valid greatly improved sensitivity with relatively little impact on specificity. For error scores, convergent validity was found to be adequate and divergent validity was found to be excellent. Although response times showed promise for assisting in the detection of feigned impairment, divergent and convergent validity were weaker, suggesting somewhat less utility than error scores.

Adolescent↗

Report of the Validation and Technology Transfer Committee of the Johns Hopkins Center for Alternatives to Animal Testing. Framework for validation and implementation of in vitro toxicity tests.

The development and application of in vitro alternatives designed to reduce or replace the use of animals, or to lessen the distress and discomfort of laboratory animals, is a rapidly developing trend in toxicology. However, at present there is no formal administrative process to organize, coordinate, or evaluate validation activities. A framework capable of fostering the validation of new methods is essential for the effective transfer of new technological developments from the research laboratory into practical use. This committee has identified four essential validation resources: chemical bank(s), cell and tissue banks, a data bank, and reference laboratories. The creation of a Scientific Advisory Board composed of experts in the various aspects and endpoints of toxicity testing, and representing the academic, industrial and regulatory communities, is recommended. Test validation acceptance is contingent upon broad buy-in by disparate groups in the scientific community-academics, industry and government. This is best achieved by early and frequent communication among parties and agreement upon common goals. It is hoped that the creation of a validation infrastructure composed of the elements described in this report will facilitate scientific acceptance and utilization of alternative methodologies and speed implementation of replacement, reduction and refinement alternatives in toxicity testing.

Animal Testing Alternatives↗

Validity and reliability of nursing workload measurement systems: review of validity and reliability theory.

Nursing workload measurement systems (WMSs) are used in inpatient and outpatient settings for staffing, scheduling, and budgeting. The nurse administrator can use WMS data to make wise decisions in these key areas providing the data are reliable and valid. Unfortunately, in most institutions, attention to issues of reliability and validity occurs only at system implementation and then the systems are left unattended. This article provides an overview of validity and reliability as it relates to WMSs. Part Two of this article will demonstrate how validity and reliability theory can be operationalized in an ongoing program for maintaining WMS reliability and validity.

Humans↗

Content validity, face validity and comprehensiveness of generic quality-of-life measures in adults and children with rare genetic conditions and their carers: a think aloud qualitative study.

PURPOSE: This study aims to assess the content validity, face validity and comprehensiveness of the: (a) EQ-5D-5L, EQ-HWB, and ASCOT SCT4, for adults with rare genetic conditions; (b) the EQ-5D-5L, EQ-HWB, and ASCOT-carer for carers of adults or children with rare genetic conditions; and (c) the EQ-5D-Y-5L carer proxy-complete for children with rare genetic conditions. METHODS: In total, 60 qualitative think-aloud interviews were conducted in Australia and England to understand individuals' thought process during the completion of the QoL measures. Participants were subsequently led through a semi-structured discussion. Transcripts were analysed for whether participants demonstrated understanding of the measures and thematic analysis was conducted on responses to the semi-structured discussion. RESULTS: The majority of participants showed good understanding and supported the validity of the measures for people experiencing rare conditions. For carers, however, a broader evaluative space than health-related QoL was preferred. Several non-health domains were identified as important to both patients and carers, including treatment availability, impact on employment and finance, information and uncertainty, medication and carer burden, impact of passing on a condition, relationships and social connection, and experience with the healthcare system. CONCLUSION: This study provides some support for the face validity and comprehensiveness of the measures for people experiencing rare conditions. However, several participants felt that the narrow health domains were inadequate to capture the breadth of their lived experience. Future research should explore the extent to which the measures capture differences and changes in the QoL domains identified as important to patients and carers.

Humans↗

Clinical validity of the Mattis Dementia Rating Scale in detecting Dementia of the Alzheimer type. A double cross-validation and application to a community-dwelling sample.

OBJECTIVE: To assess the clinical validity of the Dementia Rating Scale (DRS) in detecting patients with dementia of the Alzheimer type (DAT). BACKGROUND: The DRS is widely used to evaluate cognitive functioning in older adults. Adequate normative data are unavailable; studies addressing the clinical validity of the DRS are limited by small sample sizes. DESIGN AND METHODS: Administered the DRS to 254 outpatients with DAT and 105 healthy elderly subjects. Performed (1) multiple regressions of demographic factors on the DRS and its subscales; (2) derivation of optimal DRS cutoff scores using receiver operating characteristic curves; (3) double cross-validation with stepwise logistic regressions; and (4) application of results to a community-dwelling sample. RESULTS: Age- and education-adjusted DRS scores were computed. The optimal DRS cutoff score for DAT of 129 or less revealed a sensitivity of 98% and a specificity of 97%. The logistic regressions resulted in a combination of the Memory and Initiation/Perseveration subscales that correctly classified 98% of all subjects, 92% of a subsample of 76 patients with mild DAT, and 100% of the 51 patients with autopsy-confirmed DAT. The resultant equation was then applied to a community-dwelling sample (238 healthy elderly subjects and 44 patients with DAT): 91% of patients and 93% of normal subjects were correctly classified. Of an additional 77 individuals with questionable DAT, 43 were classified as demented and 34 were classified as nondemented. CONCLUSIONS: The DRS is a clinically valid psychometric test for the detection of DAT. The Memory and Initiation/Perseveration subscales are its best discriminative indexes for an abbreviated version.

Aged↗

Verbal Concept Attainment Test: cross-validation and validation of a booklet form.

Conducted this study to cross-validate the Verbal Concept Attainment Test as a measure of potential value in neuropsychological assessment and to validate a booklet form of this test. Two samples of 75 patients referred for neuropsychological examination were studied. In both samples the pattern of relationship between the VCAT and a number of widely used neuropsychological measures closely paralleled the pattern reported in the initial validation study. The pattern of relationships with the booklet form was also very similar to the pattern of relationships between the neuropsychological measures and the Impairment Index from the Halstead-Reitan Battery. It was concluded that these data provided evidence of the stability of this test across samples and that the booklet form appeared to be an equally valid measure.

Adolescent↗

Who checks the checkers? Four validation tools applied to eight atomic resolution structures. EU 3-D Validation Network.

Eight protein crystal structures, which have been refined against X-ray diffraction data extending to atomic resolution, 1.2 A or better, were inspected using four different validation tools, PROCHECK, PROVE, SQUID and WHATCHECK. Two general questions were addressed. (1) Do the structures imply changes in "expected" stereochemical properties and are the target values used for restraints in the validation programs and the refinement protocol appropriate? (2) Can errors in models be detected and how reliable are the coordinates after refinement? Preliminary analysis by members of the network led to modifications both to the validation programs and to the refinement protocols. The results of the final analyses are reported here. Apparent discrepancies in cell dimensions were identified. Most stereochemical properties are shown to be more tightly clustered than for lower resolution analyses. In contrast the omega angle has a wider distribution. The validation software is generally available and can be accessed at servers listed at the end of the paper.

Bacterial Proteins↗

Relative validity of a food frequency questionnaire among tin miners in China: 1992/93 and 1995/96 diet validation studies.

OBJECTIVE: Diet validation research was conducted to compare the respondents' reporting of dietary intake in a food frequency questionnaire (FFQ) with intake reported in food recalls. Because the population received annual salary increments that could modify food intake, diet validation studies (DVSs) were conducted during two time intervals. DESIGN: A 99-item FFQ was administered by an interviewer twice in a 1-year interval, and responses to each FFQ item were compared with 28 days of interviewer-administered food recalls that were collected in four 1-week intervals during each season of 1992/93. The second validation study in 1995/96 had a similar design to the earlier one. SETTING: A prospective cohort study of lung cancer among tin miners in China was initiated in 1992, with dietary and other risk factors updated annually. SUBJECTS: Among a cohort of high risk tin miners for lung cancer, two different samples (n = 141 in 1992/93, and n = 113 in 1995/96) for each diet validation study were randomly selected from four mine units, that were representative of all worker units. RESULTS: Miners reported a significantly higher average frequency of intake of foods in the food recalls than the FFQ, with few exceptions. Deattenuated Pearson correlation coefficients of the frequency of food intake between the FFQ and food recalls were in the range of -0.40 to 0.72 in both studies, with higher positive correlations for beverages and cereal staples than for animal protein sources, vegetables, fruits and legumes. The percentage of individuals with exact agreement in the extreme quartiles of intake in the food recalls and FFQ ranged from 0 to 100% in both studies. CONCLUSIONS: Among Chinese miners, the range in correlations between the food recalls and the FFQ were due to: (i) market availability of foods during the food recall weeks compared to their annual reported intake in the FFQ; (ii) cultural perception of time; and (iii) differences in how the intake of mixed dishes and their multi-ingredient foods were reported in the recalls vs. the FFQ. The range in the percentage of agreement in the same quartiles and the changes in food intake over time may have implications for the analysis of the diet-disease relationship in this cohort.

Analysis of Variance↗

Development and validation of the Validity Indicator Profile.

The Validity Indicator Profile (VIP; Frederick, 1997) is a two-alternative forced choice (2AFC) procedure intended to identify when the results of cognitive and neuropsychological testing may be invalid because of malingering or other problematic response styles. The test consists of 100 problems that assess nonverbal abstraction capacity and 78 word-definition problems. The VIP attempts to establish whether an individual's performance in an assessment battery should be considered representative of his or her true overall capacities (valid or invalid). Performances classified as valid are classified as "compliant" and reflect a high effort to respond correctly. Performances classified as invalid are subclassified as "careless" (low effort to respond correctly), "irrelevant" (low effort to respond incorrectly), or "malingering" (high effort to respond incorrectly). The VIP development sample included 944 nonclinical participants and 104 adults undergoing neuropsychological evaluation. The cross-validation sample consisted of 152 nonclinical participants, 61 brain-injured adults, 49 individuals considered to be at risk for malingering, and 100 randomly generated VIP protocols. The nonverbal subtest of the VIP demonstrated an overall classification rate of 79.8%, with 73.5% sensitivity and 85.7% specificity. The verbal subtest of the VIP demonstrated an overall classification rate of 75.5%, with 67.3% sensitivity and 83.1% specificity.

Adolescent↗