PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Internal validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Psychiatric morbidity in primary public health care: a Nordic multicentre investigation. Part I: method and prevalence of psychiatric morbidity.

The prevalence of mental illness in five different Scandinavian primary care populations was investigated in this study. Patients consecutively consulting their general practitioner a particular week-day were included in the study. Initially the SCL-25 was applied and next the high scores and a sample of the low scores were interviewed by the PSE. In the analysis the screening procedure was first validated. The internal validity of the SCL was tested by means of Rasch latent structure analysis and the external validity tested by ROC/QROC analysis. Based on this, a short 8-item version of the SCL was developed. The prevalence of mental illness in all centres was 0.26 with a minimum of 0.14 in Nacka and a maximum of 0.34 in Turku.

Adolescent↗

Evading detection on the MMPI-2: does caution produce more realistic patterns of responding?

Studies on MMPI and MMPI-2 malingering indexes often sacrifice generalizability in an attempt to control internal validity. This study improves external validity while still maintaining internal validity by providing graduate student participants with a realistic context for malingering on the MMPI-2 (n=94) and MMPI (n=30). Contextual parameters include a realistic life predicament, psychological knowledge, an incentive, the presence versus absence of a specific diagnosis, and a caution to be realistic. This study found that cautioning participants not to overexaggerate their responses significantly improves their ability to evade detection on the MMPI-2 and MMPI. Standard malingering indexes (Infrequency, F; Back Side, F, Fb; F-Correction, F-K; and Infrequency-Psychopathology, F(p)) were insufficiently sensitive in identifying simulators using common cutoff scores for these cautious simulators.

Adult↗

The Bech-Rafaelsen Mania Scale in clinical trials of therapies for bipolar disorder: a 20-year review of its use as an outcome measure.

Over the last two decades the Bech-Rafaelsen Mania Scale (MAS) has been used extensively in trials that have assessed the efficacy of treatments for bipolar disorder. The extent of its use makes it possible to evaluate the psychometric properties of the scale according to the principles of internal validity, reliability, and external validity. Studies of the internal validity of the MAS have demonstrated that the simple sum of the 11 items of the scale is a sufficient statistic for the assessment of the severity of manic states. Both factor analysis and latent structure analysis (the Rasch analysis) have been used to demonstrate this. The total score of the MAS has been standardised such that scores below 15 indicate hypomania, scores around 20 indicate moderate mania, and scores around 28 indicate severe mania. The inter-observer reliability has been found to be high in a number of studies conducted in various countries. The MAS has shown an acceptable external validity, in terms of both sensitivity and responsiveness. Thus, the MAS was found to be superior to the Clinical Global Impression scale with regard to responsiveness, and sensitivity has been found to be adequate, with the MAS able to demonstrate large drug-placebo differences. Based on pretreatment scores, trials of antimanic therapies can be classified into: (i) ultrashort (1 week) therapy of severe mania; (ii) short-term therapy (3 to 8 weeks) of moderate mania; (iii) short-term therapy of hypomanic or mixed bipolar states; and (iv) long-term (12 months) therapy of bipolar states. The responsiveness of MAS is such that the scale has been able to demonstrated that typical antipsychotics are effective as an ultrashort therapy of severe mania; that lithium and anticonvulsants are effective in the short-term therapy of moderate mania; and that atypical antipsychotics, electroconvulsive therapy (ECT) and transcranial magnetic stimulation seem to have promising effects in the short-term therapy of moderate mania. In contrast, the scale has been used to demonstrate that calcium antagonists (e.g. verapamil) are ineffective in the treatment of mania. MAS has also been used to add to the literature on the evidence-based effect of lithium as a short-term therapy for hypomania or mixed bipolar states and as a long-term therapy of bipolar states.

Antimanic Agents↗

A method for shortening instruments using the Rasch model. Validation on a hand functional measure.

BACKGROUND: Many measurement instruments, particularly measures of hand functional ability, frequently comprise a large number of items. Reduced versions of these instruments can facilitate their use. This work proposes a new method for shortening an instrument. METHODS: The method proposed was based on a scale of item difficulty calculated using the Rasch model. It was applied on a hand functional measure comprising 67 tests. The sample included 194 patients with hand lesions. The shortened instrument obtained was compared with those provided by classic methods used in the literature, with item random choice, and with shortened versions proposed by four independent experts, two rehabilitation physicians and two occupational therapists, who are clinicians familiar with the tool. All the statistical analyses were carried out on a random sub-group of two-thirds of the sample. A cross validation was then carried out on the remaining third. RESULTS: The reduction obtained had score non significantly different from that of the original instrument. In addition, the intra-class correlation coefficient and the Cronbach alpha coefficient were high. Among the different degrees of reduction investigated, the 12-item version seemed to be appropriate. Our method appeared to provide better results in terms of discriminant validity and internal validity than the choices of the four experts. The reductions produced were also better than those obtained by classic methods based on principal component analysis and multiple linear regression, as well as those obtained by random choices of items. CONCLUSION: The method presented is pertinent and useful. The reduction obtained appeared to be better than the choices of experts and the reductions provided by classic methods. The method could be used in other fields.

Accidents↗

Development and validation of a French obesity-specific quality of life questionnaire: Quality of Life, Obesity and Dietetics (QOLOD) rating scale.

OBJECTIVE: To develop and validate a new health related quality of life (HRQOL) questionnaire specific to obesity and its management. METHODS: This study was in two parts. The first (Study 1) consisted of the creation of a new tool derived from the American "Impact of Weight on Quality of Life Questionnaire" (IWQOL, 74 items) by adding to it a 17 items specific complementary module. This initial questionnaire (91 items) was reduced so as to obtain a questionnaire adapted to socio-cultural factors of obesity and dietary weight management in France. The objective of the second (Study 2) was to validate this final questionnaire by evaluating its psychometric properties: construction validity, internal reliability, concurrent validity in relation to a generic questionnaire, the SF-12, clinical validity by studying the effects of age, gender and body mass index (BMI), and reproducibility. RESULTS: The results of Study 1, obtained in 128 obese patients (mean age: 42.5 12.1, BMI: 34.5 2.8 kg/m2, women: 83.6%) enabled reduction of the 91 questionnaire items to 36, grouped into 5 dimensions: physical impact, psycho-social impact, sex life, comfort with food and diet experience. Two hundred and twelve patients (mean age: 43.3 12.2, BMI: 35.8 7.4 kg/m2, women: 77.7%) were included in Study 2, among whom 75 filled out the questionnaire twice at a one week interval. Analyses enabled verification of the construction validity and internal reliability (Cronbach alpha > 0.7) of the questionnaire as well as its concurrent validity in relation to summarized SF-12 scores and its clinical validity. The "physical impact" dimension was significantly influenced by BMI and age, the dimensions "sex life" and "diet experience" by the factors gender and BMI, while "psycho-social impact" was influenced by the 3 factors cited. Its reproducibility was also deemed satisfactory (intra-class correlation coefficient > 0.8). CONCLUSION: This new questionnaire, called the "Echelle Qualité de Vie, Obésité et Diététique (EQVOD)"/"Quality of Life, Obesity and Dietetics (QOLOD)" rating scale is sufficiently reliable and reproducible to be used in clinical practice. It is a simple tool adapted to socio-cultural factors of obesity in France, enabling taking into account of the effects of dietary management on the HRQOL of obese people.

Activities of Daily Living↗

Walking index for spinal cord injury (WISCI): an international multicenter validity and reliability study.

STUDY DESIGN: Construction of an international walking scale by a modified Delphi technique. OBJECTIVE: The purpose of the study was to develop a more precise walking scale for use in clinical trials of subjects with spinal cord injury (SCI) and to determine its validity and reliability. SETTING: Eight SCI centers in Australia, Brazil, Canada (2), Korea, Italy, the UK and the US. METHODS: Original items were constructed by experts at two SCI centers (Italy and the US) and blindly ranked in an hierarchical order (pilot data). These items were compared to the Functional Independence Measure (FIM) for concurrent validity. Subsequent independent blind rank ordering of items was completed at all eight centers (24 individuals and eight teams). Final consensus on rank ordering was reached during an international meeting (face validation). A videotape comprised of 40 clips of patients walking was forwarded to all eight centers and inter-rater reliability data collected. RESULTS: Kendall coefficient of concordance for the pilot data was significant (W=0. 843, P<0.001) indicating agreement among the experts in rank ordering of original items. FIM comparison (Spearman's rank correlation coefficient=0.765, P<0.001) showed a theoretical relationship, however a practical difference in what is measured by each scale. Kendall coefficient of concordance for the international blind hierarchical ranking showed significance (W=0.860, P<0.001) indicating agreement in rank ordering across all eight centers. Group consensus meeting resulted in a 19 item hierarchical rank ordered 'Walking Index for Spinal Cord Injury (WISCI)'. Inter-rater reliability scoring of the 40 video clips showed 100% agreement. CONCLUSIONS: This is the first time a walking scale for SCI of this complexity has been developed and judged by an international group of experts. The WISCI showed good validity and reliability, but needs to be assessed in clinical settings for responsiveness.

Australia↗

Validation of the postimplantation rat whole-embryo culture test in the international ECVAM validation study on three in vitro embryotoxicity tests.

A detailed report is presented on the performance of the postimplantation rat whole-embryo culture (WEC) test in a European Centre for the Validation of Alternative Methods (ECVAM)-sponsored formal validation study on three in vitro tests for embryotoxicity. Twenty coded test chemicals, classified as non-embryotoxic, weakly embryotoxic or strongly embryotoxic on the basis of their in vivo effects in animals and/or humans, were tested in four laboratories. The outcome showed that the WEC test can be considered to be a scientifically validated test, which is ready for consideration for use in assessing the embryotoxic potentials of chemicals for regulatory purposes.

Animal Testing Alternatives↗

Validation of the Bech-Rafaelsen Mania Scale using latent structure analysis.

The essential criteria of internal validity have not been sufficiently evaluated for any mania rating scale, although the fulfillment of such criteria is a prerequisite for summing the item scores to give a total score reflecting the severity of mania, and for comparing total scores across patient groups that differ with regard to variables such as age and sex. This study investigated the internal validity of the Bech-Rafaelsen Mania Scale (MAS), based on the ratings of 100 consecutively admitted drug-free DSM-III-R manic patients. Application of logistic latent structure models did not statistically confirm the additivity of the MAS. However, a modified MAS (the MAS-M) arising from the analyses fulfilled the measurement model. Transferability of the MAS-M across age and sex was also confirmed. The MAS-M showed an acceptable concurrent validity and an adequate sensitivity in discriminating between responders and non-responders among patients participating in a drug trial. The MAS-M presented here is the first mania rating scale that has been shown to fulfil statistical criteria for internal validity.

Adult↗

External validity: the neglected dimension in evidence ranking.

Evidence that is both accurate (internally valid) and relevant (externally valid) is needed to decide which treatment is best for a particular patient. Evidence rankings facilitate the marshalling of evidence on clinical decisions in the common context of an overwhelming number of studies, some with conflicting results. Evidence from randomized control trials is typically ranked above evidence from non-experimental studies since rankings are based primarily, if not exclusively, on considerations of internal validity. We propose that evidence rankings should consider equally both internal and external validity. External validity includes how closely the study population, the institution types in the study, the types of physicians in the study, the role of clinician decision-making (e.g. dose adjustment) in the study, and the role of patient preferences in the study resemble those in actual practice. The example of spironolactone use in heart failure illustrates the danger in using evidence that is internally but not externally valid. Ideally, a treatment should only be used when both internally and externally valid evidence indicates that it will be useful for the particular patient.

Case-Control Studies↗

Data collection methods in prospective economic evaluations: how accurate are the results?

OBJECTIVES: Often in economic evaluations a division is made between those studies that have a high level of accuracy versus those that are easily generalized. This interstudy dichotomy is often translated into prospective, randomized controlled trials with high internal validity and observational and modeling studies with a high level of external validity. This article challenges this conventional view and examines intrastudy effects on validity. METHOD: A review and summary of the literature was conducted in order to assess the impact that data collection strategies will have on internal validity. Two scenario models were created in order to gain a preliminary understanding of the magnitude of the problem. RESULTS: Data collection strategies have an impact on the level of internal validity found in an economic evaluation. Comparisons of studies that are prospective in nature is misleading as data collection strategy can lead to different resource and cost estimates even when all other relevant factors are similar. It is possible to shift and improve the level of validity by combining different collection methods. CONCLUSIONS: Instead of viewing internal and external validity as polar opposites, validity should be considered in terms of a continuum within a particular study. The use of proxies to collect resource utilization estimates, the reliance on patient self-reported data, and the method of collecting this type of data all impact the validity of study results. National guidelines for the economic evaluation of agents and devices should consider this issue in more depth, and existing evidence rankings should be adapted to be more appropriate to pharmacoeconomic studies.

Journal Article↗

Validating the International Classification of Functioning, Disability and Health Comprehensive Core Set for Rheumatoid Arthritis from the patient perspective: a qualitative study.

OBJECTIVE: To validate the International Classification of Functioning, Disability and Health (ICF) Comprehensive Core Set for Rheumatoid Arthritis (RA) from the patient perspective. METHODS: Patients with RA were interviewed about their problems in daily functioning. Interviews were tape recorded and transcribed verbatim. Interview texts were divided into meaning units. The concepts contained in these meaning units were linked to the ICF according to 10 established linking rules. Of the transcribed data, 15% were analyzed and linked by a second health professional. The degree of agreement was calculated using the kappa statistic. RESULTS: Twenty-one patients were interviewed. Two hundred twenty different concepts contained in 367 meaning units were identified in the qualitative analysis of the interviews and linked to 109 second-level ICF categories. Of the 76 second-level categories from the ICF RA Core Set, 63 (83%) were also found in the interviews. Twenty-five second-level categories, which are not part of the current ICF RA Core Set, were identified in the interviews. The result of the kappa statistic for agreement was 0.62 (95% boot-strapped confidence interval 0.59-0.66). CONCLUSION: The validity of the ICF RA Core Set was supported by the perspective of individual patients. However, some additional issues raised in this study but not covered in the current ICF RA Core Set need to be investigated.

Adolescent↗

DSM-IV internal construct validity: when a taxonomy meets data.

The use of DSM-IV based questionnaires in child psychopathology is on the increase. The internal construct validity of a DSM-IV based model of ADHD, CD, ODD, Generalised Anxiety, and Depression was investigated in 11 samples by confirmatory factor analysis. The factorial structure of these syndrome dimensions was supported by the data. However, the model did not meet absolute standards of good model fit. Two sources of error are discussed in detail: multidimensionality of syndrome scales, and the presence of many symptoms that are diagnostically ambiguous with regard to the targeted syndrome dimension. It is argued that measurement precision may be increased by more careful operationalisation of the symptoms in the questionnaire. Additional approaches towards improved conceptualisation of DSM-IV are briefly discussed. A sharper DSM-IV model may improve the accuracy of inferences based on scale scores and provide more precise research findings with regard to relations with variables external to the taxonomy.

Adolescent↗

Complex regional pain syndrome: are the IASP diagnostic criteria valid and sufficiently comprehensive?

This is a multisite study examining the internal validity and comprehensiveness of the International Association for the Study of Pain (IASP) diagnostic criteria for Complex Regional Pain Syndrome (CRPS). A standardized sign/symptom checklist was used in patient evaluations to obtain data on CRPS-related signs and symptoms in a series of 123 patients meeting IASP criteria for CRPS. Principal components factor analysis (PCA) was used to detect statistical groupings of signs/symptoms (factors). CRPS signs and symptoms grouped together statistically in a manner somewhat different than in current IASP/CRPS criteria. As in current criteria, a separate pain/sensation criterion was supported. However, unlike in current criteria, PCA indicated that vasomotor symptoms form a factor distinct from a sudomotor/edema factor. Changes in range of motion, motor dysfunction, and trophic changes, which are not included in the IASP criteria, formed a distinct fourth factor. Scores on the pain/sensation factor correlated positively with pain duration (P<0. 001), but there was a negative correlation between the sudomotor/edema factor scores and pain duration (P<0.05). The motor/trophic factor predicted positive responses to sympathetic block (P<0.05). These results suggest that the internal validity of the IASP/CRPS criteria could be improved by separating vasomotor signs/symptoms (e.g. temperature and skin color asymmetry) from those reflecting sudomotor dysfunction (e.g. sweating changes) and edema. Results also indicate motor and trophic changes may be an important and distinct component of CRPS which is not currently incorporated in the IASP criteria. An experimental revision of CRPS diagnostic criteria for research purposes is proposed. Implications for diagnostic sensitivity and specificity are discussed.

Adult↗

The effects of peer review and evidence quality on judge evaluations of psychological science: are judges effective gatekeepers?

Scientifically trained and untrained judges read descriptions of an expert's research in which the peer review status and internal validity were manipulated. Seventeen percent of the judges said they would admit the expert evidence, irrespective of its internal validity. Publication in a peer-reviewed journal also had no effect on judges' decisions. Training interacted with the internal validity manipulation. Scientifically trained judges rated valid evidence more positively than did untrained judges. Untrained judges rated a study with a confound more positively than did trained judges. Training did not affect judge evaluations of studies with a missing control group or potential experimenter bias. Admissibility decisions were correlated with judges' perceptions of the study's validity, jurors' ability to evaluate scientific evidence, and the effectiveness of cross-examination and opposing experts to highlight flaws in scientific methodology.

Adult↗

A scale for home visiting nurses to identify risks of physical abuse and neglect among mothers with newborn infants.

OBJECTIVE: The aim was to construct and test the reliability (utility, internal consistency, interrater agreement) and the validity (internal validity, concurrent validity) of a scale for home visiting social nurses to identify risks of physical abuse and neglect in mothers with a newborn child. METHOD: A 71-item scale was constructed based on a literature review and focus group sessions with social nurses and paraprofessionals who had experience with underprivileged families. This scale was applied in a random sample of 40 home visiting social nurses, who collected data in a sample of 373 nonabusive and 18 abusive/neglectful mothers with a newborn child. RESULTS: Items with prevalence rates below 5% and items making no significant difference between maltreating and non-maltreating mothers were omitted. The final version contained 20 items. This scale showed high internal consistency (alpha = .92) and high interrater reliability (r = .97). Exploratory factor analysis yielded a three-factor solution: Isolation (8 items, explaining 62.17% of the common variance), Psychological complexity (6 items, 18.86%), and Communication problems (6 items, 8.41%). Scores on Communication problems and Isolation significantly predicted scores on a social deprivation scale, which significantly distinguished maltreating from non-maltreating mothers. Mothers scoring high on Communication problems or Isolation obtained higher scores for social deprivation than low-scoring mothers. CONCLUSIONS: Home visiting nurses can identify risks for physical abuse and neglect among mothers with a newborn infant by focusing on signs of social isolation, distorted communication and psychological problems.

Adult↗

[Study on the reliability and validity of international physical activity questionnaire (Chinese Vision, IPAQ)].

OBJECTIVE: To study the reliability and validity of Chinese version of International Physical Activity Questionnaire (IPAQ) and to provide an instrument for physical activity measurement in Chinese-spoken population. METHODS: Test-retest reliability was systemically assessed in 94 participants sampled from college students. Questionnaires were completed twice with a three-day interval. The validity was established in 39 volunteers by Caltrac accelerometer monitoring and 24-hour activity recording for seven consecutive days. RESULTS: Both long vision (LV) and short vision (SV) had intraclass correlation coefficients above 0.7 for physical activity. The total energy expenditure measured by LV, SV and PA records were 264.5 +/- 260.9, 185.4 +/- 128.9 (compared with activity records, P < 0.05) and 250.5 +/- 141.2 MET-min/d respectively. Energy expenditure of moderate physical activity were 81.7 +/- 165.4, 32.0 +/- 42.5 (compared with activity record, P < 0.05) and 61.3 +/- 72.0 MET-min/d. Caltrac accelerometer was moderately correlated with LV (r = 0.50) and SV (r = 0.63) while SV measured total daily energy expenditure was lower than activity records. When participants were categorized into two groups according to their time spent in physical activity above or below the target level, proportions of agreement of questionnaires and 24-hour activity records were high, including vigorous physical activity above 90% and moderate physical activity above 70%. LV, SV and activity records were measured during sedentary condition at an approximate level. CONCLUSIONS: Both LV and SV of IPAQ appeared to have acceptable reliability and validity, compared to other physical activity instruments that were used in various large epidemiological studies. The total or physical energy expenditures were similar between LV and activity records. For activity levels, the proportion of agreement were similar between activity records and LV or SV. However, SV underestimated the energy expenditure of total and moderate physical activity.

Adult↗

Estimating sample size in clinical studies: basic methodological principles.

In order to be valid, clinical studies must be methodologically rigorous. The internal validity of a study is of crucial importance: a study is valid if its results are an unbiased estimation of the true result. In this case, the validity is internal because it refers to the group of patients under study and not necessarily different ones (external validity or applicability). Internal validity in clinical research is achieved through rigorous design, data collection and appropriate analysis, and is threatened by bias (systematic errors) or chance (random variation of the phenomena under study). Regardless of the type of study (analytic, descriptive, etc.), the characteristics of its sample are fundamental for the validity of the results. The sampling methods are crucial if the study patients are to be representative of the population to which one desires to extrapolate the results. One of the most fundamental characteristics of a sample is its size. Even the best executed study may fail to answer the research question if the sample size is too small. On the other hand, a study with too large a sample is harder to conduct and more costly. The goal of planning the sample size is to estimate the appropriate number of research subjects for the study. In this paper we will present and discuss the methodological principles underlying calculation of sample size: outcomes, type I and II error, alpha and beta, study power and variability.

Clinical Trials as Topic↗

Predictive validity and internal consistency of the pre-hospital index measured on-site by physicians.

Physiological measures of injury are used as triage tools to identify patients that require treatment in trauma centres. The Pre-Hospital Index (PHI) is based on systolic blood pressure, pulse, respiratory rate, (level of) consciousness, and presence of penetrating injury. The present study evaluated the validity and internal consistency of the PHI. The study was based on 628 patients assessed by physicians at the scene. Mean age was 38.7 years (SD = 24.8), and 65% were male. Motor vehicle collisions caused the injury for 45%. The majority had head/neck (56%) and extremity (45%) injuries. Mean PHI was 4.62 (SD = 5.77), 40% had a PHI of zero, 6% between 1 and 3, 32% between 4 and 7, and 21% greater than 7. The associations between PHI and rates of hospital admission, surgery, ICU treatment, mortality, duration of hospitalization, and length of ICU stay were significant (p < 0.001). A total of 260 (41.4%) patients had major trauma requiring treatment at a trauma centre. A PHI > 3 had 83% sensitivity and 67% specificity for identifying these patients. Internal consistency of the PHI variables was above the acceptable limits. This study has shown that the PHI is a valid and reliable physiological measure of injury severity and field triage tool.

Accidents, Traffic↗