PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “External validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning↗

Representativeness and response rates from the Domestic/International Gastroenterology Surveillance Study (DIGEST).

BACKGROUND: The Domestic/international Gastroenterology Surveillance Study (DIGEST) examined the prevalence of upper gastrointestinal symptoms among the general population in 10 countries, and the impact of these symptoms on healthcare usage and quality of life. This report discusses the validation of the DIGEST sample and reviews the response rates from the survey. METHODS: External validation of the DIGEST sample was conducted by comparing the age, age by gender and annual household incomes of the sample with census-derived data. A comparison was also made between Psychological General Well-Being Index (PGWBI) scores from study subjects in the Scandinavian countries and the USA and the total sample population norms. RESULTS: Under- and oversampling, defined as > or =5% difference from the population norms, was evident in eight out of 10 countries, but no systematic bias was evident. The final distribution of the sample by gender was 51% female and 49% male. Although differences in PGWBI scores were noted between DIGEST subjects and population norms, these differences were <0.30 standard deviations--markedly below the difference considered as relevant for the PGWBI. Response for the survey in individual countries ranged from 17% in the USA to 61% in Norway, with a survey-wide rate of 27%. The overall response rate, including primary non-respondents, was 13.4%. The majority of nonresponse (51.4%) was attributed to failure to establish contact with the subjects, with 41.7% of subjects declining to be interviewed and the remaining 6.9% of subjects not meeting the age and sex criteria used for the survey. CONCLUSIONS: The DIGEST sample exhibited good external validity, providing a foundation for comparison between data derived from individual countries in the survey.

Adult↗

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n&#xa0;=&#xa0;907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n&#xa0;=&#xa0;35), colorectal cancer (n&#xa0;=&#xa0;21), and pancreatic cancer (n&#xa0;=&#xa0;9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans↗

How generalizable are the effects of smoking prevention programs? Refusal skills training and parent messages in a teacher-administered program.

This study investigated both substantive and methodological issues associated with school-based smoking prevention programs. Substantive issues included the efficacy of a refusal skills training curriculum and of parent messages mailed to students' homes. Methodological issues included the effects of assigning classrooms versus entire schools to experimental conditions and determination of the effects of attrition on internal and external validity. Results revealed differential impact for different subgroups of adolescents. The refusal skills program produced lower rates of smoking than the control condition for students who were smokers at the pretreatment assessment but may have produced detrimental effects among males who were nonsmokers at pretest. The provision of parent messages did not affect outcome. Method of assignment (schools versus classrooms) failed to produce significant effects, and attrition did not affect internal validity. However, the above differential findings, as well as the impact of attrition on external validity, raise questions concerning the generalizability of smoking prevention programs.

Adolescent↗

Design issues for conducting cost-effectiveness analyses alongside clinical trials.

In response to rising demands for timely economic data on new medical technologies, cost-effectiveness studies are increasingly being conducted alongside clinical trials. Because of the historical differences in perspective and methods between cost-effectiveness studies and clinical trials, the design phase of these hybrid trials requires special consideration. Cost-effectiveness studies require more comprehensive evaluations of outcomes than the endpoints typically measured in clinical trials. Often, these comprehensive outcome measures (such as quality of life) prove useful for interpreting the other endpoints measured in the trial, as well as for estimating the cost-effectiveness of the intervention. In this manuscript, we discuss several aspects related to the design of joint clinical/economic trials, including study perspective, hypothesis testing, sample size estimation, and methods for collecting cost and outcome data. We also discuss issues that may limit the external validity of the cost-effectiveness results of these trials. Many potential threats to external validity can be successfully addressed if they are identified and accounted for in the design phase of the study.

Clinical Trials as Topic↗

[Underlying cause of death from external causes: validation of official data in Recife, Pernambuco, Brazil].

OBJECTIVE: To validate the underlying cause of death recorded on the death certificates for individuals under 20 years of age who died from external causes in 1995 in Recife, Pernambuco, Brazil. METHODS: We divided the study into two stages, coding and validation. In both stages we compared the official data concerning causes of death to the data we obtained during our study. We grouped the death certificates into 5 broad categories according to the cause of death; we later subdivided them into 14 categories. We also individually compared the death certificates applying the four-digit system of the International Classification of Diseases, Ninth Revision (ICD-9). We assessed the agreement between the official data and our data in terms of sensitivity and the kappa coefficient. We took as the standard the categorization of the cause of death that we had made during our investigation. RESULTS: In the coding stage, considering all the external causes of death, the overall agreement between the official data and our study data was 94% for the 5 categories, 92% for the 14 categories, and 81% for the four-digit ICD-9 system. In the validation stage the overall agreement was 94% for the 5 categories, 91% for the 14 categories, and 73% for the four-digit ICD-9 system. CONCLUSIONS: Our results suggest that for the death certificates to be reliable, the Institute of Legal Medicine must fill them out following recommended standards. In addition, hospitals and police departments must use greater care in completing the transfer slips that accompany the bodies that are sent to the Institute. More accurate data need to be generated and disseminated for a society to better understand its patterns of violence.

Adolescent↗

[Interobserver and intraobserver variation: a problem of validity in epidemiologic studies of arterial pressure].

In carrying out blood pressure epidemiologic studies there may be different factors that can affect internal and external validity and thus eliminate the inferential process. As part of the Hypertension and Risk Factors Associated Study conducted in March 1987 in Cuajimalpa de Morelos, Mexico City, 23 nursing students were standardized on the blood pressure auscultatory method using a sound picture and measuring intraobserver and interobserver agreement through intraclass correlation coefficient. Even though initial standardization sessions showed difficulties in the use of instruments and in the reading of blood pressure levels, final K (kappa) values measuring interobserver agreement increased from 0.25 to 0.86. Omega values measuring intraobserver agreement fluctuated between 0.86 and 0.98. This epidemiologic technique is proposed in order to improve internal and external validity of blood pressure studies.

Blood Pressure↗

The WHO (Ten) Well-Being Index: validation in diabetes.

BACKGROUND: In a European trial in 8 countries, the subjective well-being of patients on alternative forms of treatment for insulin-dependent diabetes was compared using the 28-item WHO Well-Being Questionnaire, covering four dimensions of depression, anxiety, energy and positive well-being. The objective of the analysis reported here has been to identify the items of the WHO questionnaire which belong to an overall index of negative and positive well-being. METHODS: Adult patients at 10 study centres in 8 countries who had been on insulin for at least 2 years were invited to participate in a randomised, cross-over trial to compare insulin pump treatment with injection therapy. At each phase, patients completed questions on well-being and general health. Internal validity of the well-being index was evaluated by Cronbach's alpha and Loevinger's and Mokken's homogeneity coefficients, as well as factor analysis. External validity was evaluated by comparisons with results of the general assessment questions and by the ability to discriminate between the alternative forms of treatment. RESULTS: 358 patients had sufficient data for analysis. Ten items were found to constitute a valid index of well-being with respect to internal and external validity. Coefficients of homogeneity were acceptable and there was evidence for both concurrent and discriminant validity. CONCLUSIONS: The WHO (Ten) well-being index includes negative and positive aspects of well-being in a single uni-dimensional scale. Its advantage lies in its ability to show overall change along the continuum of well-being, thus facilitating comparisons between patient groups and treatments. It is not specific to diabetes, and therefore may be useful as a disease-independent index of well-being in a broad range of health care studies.

Adaptation, Psychological↗

Contemporary identification of patients at high risk of early prostate cancer recurrence after radical retropubic prostatectomy.

OBJECTIVES: To develop a model that will identify a contemporary cohort of patients at high risk of early prostate cancer recurrence (greater than 50% at 36 months) after radical retropubic prostatectomy for clinically localized disease. Data from this model will provide important information for patient selection and the design of prospective randomized trials of adjuvant therapies. METHODS: Proportional hazards regression analysis was applied to two patient cohorts to develop and cross-validate a multifactorial predictive model to identify men with the highest risk of early prostate cancer recurrence. The model and validation cohorts contained 904 and 901 men, respectively, who underwent radical retropubic prostatectomy at Johns Hopkins Hospital. This model was then externally validated using a cohort of patients from the Mayo Clinic. RESULTS: A model for weighted risk of recurrence was developed: R(W)'=lymph node involvement (0/1)x1.43+surgical margin status (0/1)x1.15+modified Gleason score (0 to 4)x0.71+seminal vesicle involvement (0/1)x0.51. Men with an R(W)' greater than 2.84 (9%) demonstrated a 50% biochemical recurrence rate (prostrate-specific antigen level greater than 0.2 ng/mL) at 3 years and thus were placed in the high-risk group. Kaplan-Meier analyses of biochemical recurrence-free survival demonstrated rapid deviation of the curves based on the R(W)'. This model was cross-validated in the second group of patients and performed with similar results. Furthermore, similar trends were apparent when the model was externally validated on patients treated at the Mayo Clinic. CONCLUSIONS: We have developed a multivariate Cox proportional hazards model that successfully stratifies patients on the basis of their risk of early prostate cancer recurrence.

Adult↗

Validation of the Hamilton Depression Rating Scale and Montgommery and Asberg Rating Scales in terms of AGECAT depression cases.

OBJECTIVE: To validate the Hamilton Depression (17) and Montgommery and Asberg Depression Scales as research instruments in older depressed community residents. DESIGN: External validation against GMS/AGECAT case level in the recruitment of older community residents for an antidepressant trial. ANALYSES: Receiver operator curves were generated for each rating scale, using GMS/AGECAT case level in external criterion. The sensitivity, specificity, positive and negative predictive values of both rating instruments were examined in the whole sample and age and gender subgroups. MADRS and HAM-D cut-off scores differentiating GMS/AGECAT cases from subcases were identified. RESULTS: HAM-D cut-off score of 16 and MADRS score of 21 were identified as differentiating case from sub-case. Diagnostic accuracy of both instruments was good, reflecting good sensitivity and specificity across both genders and sub-age groups. CONCLUSIONS: Both scales performed well in this population. These scores provide researchers with externally validated and clinically relevant cut-off scores in designing trials in the management of older depressed community residents.

Aged↗

[Clinical usefulness of oligoclonal bands].

The presence of oligoclonal bands (OCB) of immunoglobulin G (IgG) is in our days the most useful finding in the study of the CSF for the diagnosis of multiple sclerosis (MS). The most sensitive method for the detection of OCB is the isoelectric focusing followed by immunoblotting. The prevalence of OCB changes in different populations with a rank of results from 60 to 95 97%. We have determined the prevalence of OCB in our population and the sensitivity and the specificity of the technique used in our laboratory. We have included 391 patients in whom we analysed the presence of OCB, subdivided in; Group 0: Diagnosed of MS, group 1: First episode of demyelinating process, group 2: Neurological disorders considered noninflammatory or nonautoimmune (NINA),group 3: Neurological disorders considered inflammatory, infectious or autoimmune (IIA). The presence of OCB was searched in CSF and serum simultaneously using isoelectric focusing and immunoblotting. In order to standardize the technique we achieved and internal and external validation. Internal validation: sensitivity and specificity (using as a control group first the group NINA and after the group IA). External validation: we choose 10 pairs of CSF/serum from patients with different diagnostics and sent to a reference laboratory ( Karolinska Institute Medical School) that was blind of our results and of the diagnostics. The prevalence of OCB in each group has been: group 0 (MS): 87.7%, group 1: 54.8%, group 2 (NINA): 17.5%, group 3(IIA): 52.7%. Sensitivity: 97.7%, specificity using group NINA as control 82.5% and using group IIA 45.7%. Concordance with the reference laboratory in 9/10 determinations. We conclude that in our population the prevalence of OCB, in patients with MS, is lower than in Northern Europe. The OCB appear in may inflammatory, autoimmune diseases, their specificity for the diagnostic of MS is low.

Autoimmune Diseases↗

Artificial Intelligence for Diagnosis, Risk Stratification, and Prognosis of Neuroblastoma - A Systematic Review and Meta-Analysis.

PURPOSE: To synthesizes evidence on artificial intelligence (AI) performance in neuroblastoma (NB) diagnosis, risk stratification, prognosis, and genomic characterization. MATERIALS AND METHODS: A systematic review and meta-analysis was conducted following PRISMA 2020 guidelines (PROSPERO: CRD42024539475) across five databases. Meta-analyses used random-effects models with logit-transformed Area Under the Curve (AUCs) and cluster-robust standard errors. AI models were classified as Machine Learning Models (MLM) or Hybrid Nomograms (HN) based on their construction methodology. RESULTS: Of 3,742 articles identified, 53 were included. MLMs demonstrated higher point estimates than radiologists in differential diagnosis (AUC: 0.87 vs. 0.83), though this difference was not statistically significant and carried substantial uncertainty. HNs achieved stronger performance in risk stratification (AUC: 0.87). AI-derived nomograms (AUC: 0.9) and gene signatures (AUC: 0.8) outperformed conventional prognostic markers descriptively. Chemotherapy response prediction remained below clinical utility thresholds across all model types. Only 33.9% of models reported calibration and 24.5% underwent external validation. CONCLUSIONS: AI demonstrates proof-of-concept across multiple NB clinical domains. However, clinical adoption remains premature given persistent gaps in external validation, calibration, dataset size, and pediatric-specific model development. Future studies should test these models prospectively in multicenter pediatric cohorts, ideally through COG or SIOPEN, using shared definitions for diagnosis, risk group, treatment response, and survival outcomes.

Humans↗

Using prognostic models in clinical infertility.

The chance that a couple who have tried to conceive for 12 months will succeed without assisted conception treatment is still higher than the chance that the same couple will benefit from treatment. In this context, it is important to assess the chance that a treatment-independent or 'spontaneous' pregnancy will occur in a couple whose wish for a child is unfulfilled. Prognostic models can be useful in this assessment. In recent years, prognostic models have been published both for the occurrence of 'spontaneous' pregnancy and for pregnancy after in vitro fertilization. This article discusses the theoretical aspects of prognostic modelling and assesses whether the current prognostic models are good enough to justify their use in clinical practice. The performance of existing models for the prediction of spontaneous conception was found to be acceptable on internal as well as on external validation. However, the performance of the existing models predicting IVF outcome was found to be disappointing on the few occasions on which such external validation has been performed.

Journal Article↗

Predictive validity of the strain index in manufacturing facilities.

The Strain Index is a job analysis method for determining if workers are exposed to increased risk of developing distal upper extremity disorders. Its predictive and external validity was initially demonstrated in a pork processing plant. The purpose of this study was to evaluate its predictive validity in two manufacturing plants. While blinded to health outcomes, investigators analyzed the right and left sides of 28 single-task jobs using the Strain Index and classified them as "hazardous" or "safe" based on the Strain Index score. Subsequently, OSHA 200 logs were used to ascertain the occurrence of distal upper extremity disorders retrospectively. If at least one such disorder occurred on the right or left side during the prior three years, that side was classified as "positive." If no such disorder was reported during the prior three years, that side was classified as "negative." When comparing sides, symmetry between morbidity and hazard classification was required. When comparing jobs, such symmetry was not required. Evidence of association between the hazard classifications and the morbidity classifications for the 56 sides and the 28 jobs was evaluated using 2 x 2 contingency tables. For the sides, the association between hazard classification and morbidity classification was statistically significant with an empirical odds ratio of 73.2. The sensitivity, specificity, positive predictive value, and negative predictive value were 1.00, 0.84, 0.47, and 1.00. Similar results were noted for the jobs--the empirical odds ratio was 106.6, and the sensitivity, specificity, positive predictive value, and negative predictive value were 1.00, 0.91, 0.75, and 1.00. While these results provide additional evidence of the Strain Index's external validity and predictive validity, it should be noted that these jobs involved the performance of single tasks.

Arm Injuries↗

Predictive validity of the Strain Index in turkey processing.

The Strain Index is a job analysis method for determining if workers are exposed to increased risk of developing distal upper extremity disorders. Its predictive and external validity was initially demonstrated in a pork processing plant. The purpose of this study was to evaluate the predictive validity of the Strain Index in one turkey processing plant. While blinded to health outcomes, investigators analyzed the right and left sides of workers in 28 jobs using the Strain Index and classified them as "hazardous" or "safe" based on the Strain Index score. Subsequently, OSHA 200 logs were used to ascertain the occurrence of distal upper extremity disorders retrospectively. If at least one such disorder had occurred on the right or left side during the previous 3 years, that side was classified as "positive." If no such disorder was reported during the previous 3 years, that side was classified as "negative." When comparing sides, symmetry between morbidity and hazard classification was required. When comparing jobs, such symmetry was not required. Evidence of association between the hazard classifications and the morbidity classifications for the 56 sides and the 28 jobs was evaluated using 2 x 2 contingency tables. For the sides, the association between hazard classification and morbidity classification was statistically significant, with an odds ratio of 22.0. The sensitivity, specificity, positive predictive value, and negative predictive value were 0.86, 0.79, 0.92, and 0.65, respectively. Similar results were noted for the jobs--the odds ratio was 50.0, and the sensitivity, specificity, positive predictive value, and negative predictive value were 0.91, 0.83, 0.95, and 0.71. These results provide additional evidence of the external validity and predictive validity of the Strain Index.

Animals↗

Development and validation of the Headache Needs Assessment (HANA) survey.

OBJECTIVE: To develop and validate a brief survey of migraine-related quality-of-life issues. The Headache Needs Assessment (HANA) questionnaire was designed to assess two dimensions of the chronic impact of migraine (frequency and bothersomeness). METHODS: Seven issues related to living with migraine were posed as ratings of frequency and bothersomeness. Validation studies were performed in a Web-based survey, a clinical trial responsiveness population, and a retest reliability population. Headache characteristics (eg, frequency, severity, and treatment), demographic information, and the Headache Disability Inventory were used for external validation. RESULTS: The HANA was completed in full by 994 adults in the Web survey, with a mean total score of 77.98 +/- 40.49 (range, 7 to 175). There were no floor or ceiling effects. The HANA met the standards for validity with internal consistency reliability (Cronbach alpha =.92, eigenvalue for the single factor = 4.8, and test-retest reliability = 0.77). External validity showed a high correlation between HANA and Headache Disability Inventory total scores (0.73, P<.0001), and high correlations with disease and treatment characteristics. CONCLUSIONS: These data demonstrate the psychometric properties of the HANA. The brief questionnaire may be a useful screening tool to evaluate the impact of migraine on individuals. The two-dimensional approach to patient-reported quality of life allows individuals to weight the impact of both frequency and bothersomeness of chronic migraines on multiple aspects of daily life.

Activities of Daily Living↗

Subject attrition in prevention research.

Subject attrition threatens the internal validity of substance abuse prevention studies because differences in the rate of attrition and the substance use behavior of remaining subjects in the different conditions could account for any differences found in substance use rates. Attrition threatens the external validity of prevention studies because, to the extent that study dropouts are different from remaining subjects, the results of the study may not be generalizable to study dropouts. Analysis of these threats to the validity of prevention studies should be routinely conducted. However, studies of alcohol and drug abuse prevention have generally failed to report or analyze subject attrition. Smoking prevention studies have more frequently reported attrition, and they have recently begun to analyze the degree to which attrition may affect the internal and external validity of the study. Evidence thus far suggests that differences in attrition across conditions do occur occasionally. The evidence is substantial that study dropouts are systematically more likely to smoke, to use other substances, and to score highly on other risk-taking measures.

Alcoholism↗

Validation of the inflammatory bowel disease questionnaire in Swedish patients with ulcerative colitis.

BACKGROUND: The Inflammatory Bowel Disease Questionnaire (IBDQ) is a disease-specific health-related quality of life (HRQOL) questionnaire including four dimensions and a sum score. The aim of this study was to assess the internal and external validity, reliability, and sensitivity of a Swedish version of the IBDQ. METHODS: Three hundred consecutive patients with ulcerative colitis completed the IBDQ and three other health-related quality of life questionnaires (the Rating Form of IBD Patient Concerns (RFIPC), the Short Form-36 (SF-36) and the Psychological General Well-Being (PGWB) index). Disease activity was evaluated using a 1-week symptom diary, blood tests and rigid sigmoidoscopy. One hundred and fourteen patients filled in the questionnaire a second time, of whom 75 had been in stable remission for over 6 months and 39 had a significant clinical change in disease activity. RESULTS: Factor analysis of the 32 IBDQ items did not support the four dimensional scores. The dimensional scores had sufficient convergent validity, but low discriminative validity and homogeneity. The homogeneity was also low for the sum score. The inter-dimensional correlations were high. The concurrent validity was supported by correlations between the dimensional scores and other measures of disease activity and HRQOL. Patients in relapse scored significantly less on the sum score and the four dimensions compared to patients in remission. The test-retest correlations for the dimensional scores were 0.40-0.76. Patients with a change in disease activity during the 6-month follow-up period had a significant change in IBDQ scores not found in those who remained in remission. CONCLUSIONS: The Swedish version of the IBDQ had external validity and was shown to be a reliable and sensitive measure of HRQOL in ulcerative colitis, though there are some concerns regarding the internal validity. The use of a sum score was not supported and the questionnaire may benefit from a redivision of items into dimensions with better homogeneity and discriminative validity.

Colitis, Ulcerative↗