PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Internal validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Psychometric properties of the simplified Chinese version of the EORTC QLQ-BR53 for measuring quality of life for breast cancer patients.

A Simplified Chinese version of the EORTC QLQ-BR53 was evaluated using responses from 233 patients with breast cancer in China by assessing the construct and criterion-related validity, internal consistency and test-retest reliability, and responsiveness as measured by score changes of the scales. Internal consistency reliability measured by Cronbach's coefficient alpha is greater than 0.75 for most multi-item scales except cognitive functioning (0.41) and breast symptoms (0.71). Test-retest reliability coefficients for all domains are greater than 0.80 with the exception of physical functioning (0.65), social functioning (0.75), appetite loss (0.75), diarrhea (0.72), and body image (0.72). Correlation and factor analysis among domains and items showed good construct validity for both QLQ-C30 and QLQ-BR23. Score changes over time were observed in most domains except emotional functioning, global health status/QOL, dyspnoea, constipation, diarrhea, financial difficulties, sexual functioning, sexual enjoyment, and breast symptoms. Therefore, the Simplified Chinese version of QLQ-BR53 shows reasonable validity, reliability, and responsiveness and can be used to measure QOL for Chinese patients with breast cancer.

Adolescent↗

[Cultural adaptation and validation of questionnaires measuring satisfaction with the French health system].

Two questionnaires measuring satisfaction of the population with regard to health care offer were constructed from measures validated in the USA (the Consumer Satisfaction Survey Questionnaire or CSS, and the Visit-Specific Satisfaction questionnaire, the VSQ). This work was comprised two stages: i) translation and cultural adaptation of the American instrument to the French health care context, implicating 6 translators, users and experts; and ii) a telephone survey in the general population (n = 706) to test the psychometric qualities of the French instrument (content and internal validity). The French version, the CSS-VF comprises 9 scales: access to primary care, access to secondary care, scope for choice, health cover, communication with and competence of GPs, communication with specialists, competence of specialists, human qualities of practitioners and overall satisfaction. The VSQ-VF, which measures satisfaction with the last medical consultation is unidimensional. The results of the psychometric analyses are good overall, and endorse the use of these scales in assessment studies.

Consumer Behavior↗

Development of a Korean version of the Female Sexual Distress Scale.

INTRODUCTION: This article presents data based on the responses of more than 100 women who contributed to the development of a Korean version of the Female Sexual Distress Scale (FSDS). AIM: The FSDS was developed to measure sexually related personal distress in women. This article aims to test the usefulness and analyze factors of the 20-item version of the FSDS in a Korean female sample. METHODS: The original two-item FSDS was translated with cultural modifications. A total of 104 healthy, married women were recruited through a survey. A second survey was undertaken after 2 weeks for test-retest reliability. Validity, internal consistency reliability, and test-retest reliability were evaluated. An exploratory factor analysis was also performed. MAIN OUTCOME MEASURES: A Korean version of the FSDS. RESULTS: The test-retest coefficients of stability over a 2-week period was 0.99 (P < 0.01). The 20 items of the FSDS have good internal consistency, with an alpha of 0.96. The FSDS discriminated between women with and without sexually related distress (t = -7.34, P < 0.01). The optimal cut-off score was 20 (sensitivity 71.4%, specificity 92.2%). By principal axis factoring, the Korean version of the FSDS was found to consist of two factors. A 16-item FSDS had good internal consistency with an alpha of 0.97. The test-retest reliability was good (r = 0.99, P < 0.01). The items of the 16-item FSDS were somewhat different from the original 12-item FSDS. CONCLUSIONS: The Korean version of the FSDS (20-item) might be a useful tool for screening sexually distressed women in Korea. Instead of the 12-item version of the original FSDS, the 16-item FSDS was validated in this study. These results could reflect cultural differences between Eastern Asian and Western societies.

Adult↗

A three-metabolite microbiota-associated signature for early risk stratification of gestational diabetes mellitus.

BACKGROUND: Gestational diabetes mellitus (GDM) is associated with adverse pregnancy outcomes and long-term metabolic and cardiovascular risk. However, oral glucose tolerance testing at 24-28 gestational weeks limits early risk stratification. Gut microbiota-associated metabolites may reflect early metabolic abnormalities, including those relevant to cardiometabolic health, but robust early-pregnancy biomarkers remain limited. METHODS: We conducted a multicenter nested case-control and prospective study involving 2,693 pregnant women. Untargeted metabolomics and metagenomics were integrated to identify GDM-associated metabolites and gut microbial alterations. Three consistently dysregulated metabolites, 3-hydroxydecanoic acid, &#x3b3;-Glu-Leu, and propionic acid, were quantified by targeted LC-MS/MS. Candidate algorithms were compared using repeated 10-fold cross-validation, and a final generalized linear model was externally and prospectively validated. RESULTS: Women who later developed GDM showed an adverse early-pregnancy metabolic profile, including higher BMI, triglycerides, and platelet count. Untargeted metabolomics identified 14 persistently altered metabolites enriched in energy, oxidative stress, and amino acid metabolism pathways. Metagenomics revealed taxonomic restructuring and coordinated microbiota-metabolite associations. The three-metabolite model achieved AUCs of 0.838 (95% CI, 0.791-0.885) in training, 0.840 (95% CI, 0.769-0.911) in internal validation, 0.955 (95% CI, 0.925-0.985) and 0.917 (95% CI, 0.875-0.958) in two external cohorts, and 0.969 (95% CI, 0.937-1.000) in the prospective cohort. CONCLUSION: Early microbiota-associated metabolic dysregulation is detectable before routine GDM diagnosis. This compact three-metabolite panel may support early GDM risk stratification and provides metabolic evidence relevant to broader cardiometabolic risk assessment in pregnancy.

Humans↗

Deep learning techniques in predicting BRAF mutation status in cutaneous melanoma from histopathologic images.

AIMS: To develop and validate a deep learning framework for discriminating BRAF mutation status in cutaneous melanoma from routine H&E whole-slide images (WSIs) as a proof-of-concept complementary approach alongside molecular testing. METHODS: We built a two-stage pipeline comprising U-Net-based tumour segmentation followed by an Inception v3 classifier. In total, 272 institutional melanoma cases with confirmed BRAF status were used for model development (training and internal validation). Generalisability was assessed in an external test set of 76 cutaneous melanoma cases from the Cancer Genome Atlas (TCGA). Dermatopathologist-defined tumour-rich regions of interest were used to train and evaluate segmentation. WSIs were processed at 20&#xd7;magnification using 512&#xd7;512 tiles; slide-level mutation probabilities were obtained by averaging the predicted probabilities across all tumour-enriched tiles. RESULTS: Inception v3 achieved area under the receiver operating characteristic curve values of 0.973 (training), 0.954 (validation) and 0.915 (TCGA testing) and outperformed a ResNet50 baseline, showing stable external generalisation. Performance remained robust in advanced pathological T-category primary tumours (pT3-T4). Tumour probability heatmaps supported spatial interpretability by localising regions contributing most strongly to predicted mutation status. CONCLUSIONS: Deep learning applied to routine H&E WSIs can infer BRAF mutation status in cutaneous melanoma with consistent performance across institutional and external cohorts. Given the observed external sensitivity and negative predictive value, the model is not suitable for rule-out use or for deferring/omitting molecular testing. Any workflow integration is future work and would require prospective validation and calibration of probability outputs in real-world clinical series.

Artificial Intelligence↗

Development of the Family Nursing Practice Scale.

This article describes the development and testing of the Family Nursing Practice Scale (FNPS). This self-report questionnaire is designed to measure perceived changes in family nursing practice including attitudes toward working with families, critical appraisal of their family nursing practice and reciprocity in the nurse-family relationship. Categories were derived from a needs assessment, competence as effective application of knowledge and skill and theoretical foundations for family assessment and intervention. Psychometric testing (content, construct validity, internal consistency, and test-retest reliability) was undertaken with 140 psychiatric nurses in Hong Kong. Practice appraisal and nurse-family relationships accounted for 56.4% of the variance. Cronbach's alpha reliability coefficients were .88 and .73 for the two subscales, respectively, and .86 for the scale overall. Test-retest reliability ranged from .62 to .93 on the individual items. The results provide preliminary evidence of the reliability and validity of the FNPS. The instrument provides quantitative and qualitative evaluation components.

Adult↗

Checklist for the qualitative evaluation of clinical studies with particular focus on external validity and model validity.

BACKGROUND: It is often stated that external validity is not sufficiently considered in the assessment of clinical studies. Although tools for its evaluation have been established, there is a lack of awareness of their significance and application. In this article, a comprehensive checklist is presented addressing these relevant criteria. METHODS: The checklist was developed by listing the most commonly used assessment criteria for clinical studies. Additionally, specific lists for individual applications were included. The categories of biases of internal validity (selection, performance, attrition and detection bias) correspond to structural, treatment-related and observational differences between the test and control groups. Analogously, we have extended these categories to address external validity and model validity, regarding similarity between the study population/conditions and the general population/conditions related to structure, treatment and observation. RESULTS: A checklist is presented, in which the evaluation criteria concerning external validity and model validity are systemized and transformed into a questionnaire format. CONCLUSION: The checklist presented in this article can be applied to both planning and evaluating of clinical studies. We encourage the prospective user to modify the checklists according to the respective application and research question. The higher expenditure needed for the evaluation of clinical studies in systematic reviews is justified, particularly in the light of the influential nature of their conclusions on therapeutic decisions and the creation of clinical guidelines.

Bias↗

Nomogram for overall survival of patients with progressive metastatic prostate cancer after castration.

PURPOSE: To develop a pretreatment prognostic model for survival of patients with progressive metastatic prostate cancer after castration using parameters that are measured during routine clinical management. PATIENTS AND METHODS: Pretreatment clinical and biochemical determinants from 409 patients enrolled onto 19 consecutive therapeutic protocols from June 1989 through January 2000 were evaluated. The factors selected were age, Karnofsky performance status (KPS), hemoglobin (HGB), prostate-specific antigen (PSA), lactate dehydrogenase (LDH), alkaline phosphatase (ALK), and albumin. These factors were combined in an accelerated failure time regression model to produce a nomogram to predict median, 1-year, and 2-year survival. The nomogram was validated internally and externally using data from a multicenter randomized trial of suramin plus hydrocortisone versus hydrocortisone alone. RESULTS: The median survival of the entire group was 15.8 months (range, 0.9 to 77.8 months); 87% have died. In multivariable analysis, KPS, HGB, ALK, albumin, and LDH were significantly associated with survival (P <.05), whereas age and PSA were not. All seven factors were included in the nomogram. When applied to the external validation data set, the nomogram achieved a concordance index of 0.67. Calibration plots suggested that the nomogram was well calibrated for all predictions. CONCLUSION: A nomogram derived from pretreatment parameters that are measured on a routine basis was constructed. It can be used to predict the median, 1-year, and 2-year survival of patients with progressive castrate metastatic disease with reasonable accuracy. The information is useful to assess prognosis, guide treatment selection, and design clinical trials.

Adult↗

International Programme for Resource Use in Critical Care (IPOC)--a methodology and initial results of cost and provision in four European countries.

BACKGROUND: A standardized top-down costing method is not currently available internationally. An internally validated method developed in the UK was modified for use in critical care in different countries. Costs could then be compared using the World Health Organization's Purchasing Power Parities (WHO PPPs). METHODS: This was an observational, retrospective, cross-sectional, multicentre study set in four European countries: France, UK, Germany and Hungary. A total of 329 adult intensive care units (ICUs) participated in the study. RESULTS: The costs are reported in international dollars ($) derived from the WHO PPP programme. The results show significant differences in resource use and costs of ICUs over the four countries. On the basis of the sum of the means for the major components, the average cost per patient day in UK hospitals was $1512, in French hospitals $934, in German hospitals $726 and in Hungarian hospitals $280. CONCLUSIONS: The reasons for such differences are poorly understood but warrant further investigation. This information will allow us to better adjust our measures of international ICU costs.

Costs and Cost Analysis↗

Construct validity of the adapted Questionnaire on Resources and Stress--short form.

The concurrent validity of the adapted version of Holroyd's (1982) short form of the Questionnaire on Resources and Stress (QRS) was examined with a sample of 103 mothers of children with mild to severe disabilities. A correlation matrix was generated using the QRS adapted short form total, its seven component factors, and six criterion measures. Results revealed significant relations among the criterion measures and their target factors as well as concurrence for the internal validity of the instrument.

Adaptation, Psychological↗

Assessing student reflection in medical practice. The development of an observer-rated instrument: reliability, validity and initial experiences.

INTRODUCTION: This study describes the development of an instrument to measure the ability of medical students to reflect on their performance in medical practice. METHODS: A total of 195 Year 4 medical students attending a 9-hour clinical ethics course filled in a semi-structured questionnaire consisting of reflection-evoking case vignettes. Two independent raters scored their answers. Respondents were scored on a 10-point scale for overall reflection score and on a scale of 0-2 for the extent to which they mentioned a series of perspectives in their reflections. We analysed the distribution of scores, the internal validity and the effect of being pre-tested with an alternate form of the test on the scores. The relationships between overall reflection score and perspective score, and between overall reflection score and gender, career preference and work experience were also calculated. RESULTS: The interrater reliability was sufficient. The range of scores on overall reflection was large (1-10), with a mean reflection score of 4.5-4.7 for each case vignette. This means that only 1 or 2 perspectives were mentioned, and hardly any weighing of perspectives took place. The values over the 2 measurements were comparable and were strongly related. Women had slightly higher scores than men, as had students with work experience in health care, and students considering general practice as a career. CONCLUSIONS: Reflection in medical practice can be measured using this semistructured questionnaire built on case vignettes. The mean score allows for the measurement of improvement by future educational efforts. The wide range of individual differences allows for comparisons between groups. The differences found between groups of students were as expected and support the validity of the instrument.

Adult↗

Predictors of outcome of epilepsy surgery: multivariate analysis with validation.

PURPOSE: To identify predictors of outcome of epilepsy surgery, using the Duke experience, applying multivariate analysis and validation techniques. To compare the results of different modeling algorithms. Few previous studies have reported multivariate analysis, or validated their results. METHODS: Records of 116 patients with focal resections for intractable epilepsy from January 1, 1980 through June 30, 1989 were analyzed. Primary outcome variable was patient's condition in second postoperative year: seizure free (except auras), or not. Three predictors of biologic interest were specified a priori for confirmatory analysis. Additional predictors were considered within exploratory analysis. Logistic regression techniques were applied to assess relations with pre- and postoperative predictors. Internal validity was assessed by repeated random selection of training and validation samples, used in conjunction with bootstrap techniques. RESULTS: By using multivariate analysis, percentage of epileptic EEG activity arising from the site of resection and either imaging localization or lack of use of invasive monitoring were the only statistically significant preoperative predictors for good outcome at 2 years. Presence of seizures within 2 months of surgery was a significant postoperative predictor for a poor outcome. Adding more variables did not result in significantly improved models. Use of validation techniques reduced the degree of optimism in the predictive value of the models. CONCLUSIONS: Pooling of data from multiple institutions is needed to attain the large sample sizes needed for multivariate analysis with validation.

Adolescent↗

The Sexual Arousal and Desire Inventory (SADI): a multidimensional scale to assess subjective sexual arousal and desire.

INTRODUCTION: Sexual arousal and desire are integral parts of the human sexual response that reflect physiological, emotional, and cognitive processes. Although subjective and physiological aspects of arousal and desire tend to be experienced concurrently, their differences become apparent in certain experimental and clinical populations in which one or more of these aspects are impaired. There are few subjective scales that assess sexual arousal and desire specifically in both men and women. AIMS: (i) To develop a multidimensional, descriptor-based Sexual Arousal and Desire Inventory (SADI) to assess subjective sexual arousal and desire in men and women; (ii) to evaluate convergent and divergent validity of the SADI; and (iii) to assess whether scores on the SADI would be altered when erotic fantasy or exposure to an erotic film was used to increase subjective arousal. METHODS: Adult men (N = 195) and women (N = 195) rated 54 descriptors as they applied to their normative experience of arousal and desire on a 5-point Likert scale. Another sample of men (N = 40) and women (N = 40) completed the SADI and other measures after viewing a 3-minute female-centered erotic film or engaging in a 3-minute period of erotic fantasy. MAIN OUTCOME MEASURES: Principal components analyses derived factors that the scale descriptors loaded onto. These factors were categorized as subscales of the SADI, and gender differences in ratings and internal validity were analyzed statistically. Factors were considered subscales of the SADI, and mean ratings for each subscale were generated and related to the other scales used to assess convergent and divergent validity. These scales included the Feeling Scale, the Multiple Indicators of Subjective Sexual Arousal, the Sexual Desire Inventory, and the Attitudes Toward Erotica Questionnaire, the Beck Depression Inventory (BDI)-II, and the Beck Anxiety Inventory. RESULTS: Descriptors loaded onto four factors that accounted for 41.3% of the variance. Analysis of descriptor loadings > or = 0.30 revealed an Evaluative factor, a Physiological factor, a Motivational factor, and a Negative/Aversive factor based on the meaning of the descriptors. Men's and women's subjective experiences of sexual desire and arousal on the Physiological and Motivational factors were not significantly different, although on the Evaluative and Negative factors, statistically significant differences were found between the genders. Mean scores on the Evaluative factor were higher for men than for women, whereas mean scores on the Negative factor were higher for women than for men. Internal consistency estimates of the SADI and its subscales confirmed strong reliability. Mean scores on the Evaluative, Motivational, and Physiological subscales of the SADI were significantly higher in the fantasy condition than in the erotic clip condition. Women had significantly higher mean scores than men on the Physiological subscale in the fantasy condition. Cronbach's alpha coefficients demonstrated excellent reliability of the SADI subscales. Evidence of convergent validity between the SADI subscales and other scales that measured the same constructs was strong. Divergent validity was also confirmed between the SADI subscales and the other scales that did not measure levels of sexual arousal, desire, or affect, such as the BDI-II. CONCLUSION: The SADI is a valid and reliable research tool to evaluate both state and trait aspects of subjective sexual arousal and desire in men and women.

Adult↗

Cultural dissimilarities in general practice: development and validation of a patient's cultural background scale.

Due to increased migration physicians encounter more communication difficulties due to poor language proficiency and different culturally defined views about illness. This study aimed to develop and validate a 'patient's cultural background scale' in order to classify patients based on culturally conditioned norms instead of on ethnicity. A total of 986 patients from 38 multi-ethnic general practices were included. From a list of 36 questions, non-contributing and non-consistent questions were deleted and from the remaining questions the scale was constructed by principal component analysis. Comparing the scale with two other methods of construction assessed internal validity. Comparing the found dimensions with known dimensions from literature assessed the construct validity. Criterion validity was determined by comparing the patient's score with criteria assumed or known to have relationship with cultural background. Criterion validity was reasonably good but poor for income. A valid patient's cultural background scale was developed, for use in large-scale quantitative studies.

Adult↗

High prevalence of erectile dysfunction after renal transplantation.

BACKGROUND AND METHODS: A cross-sectional study of multifaceted male sexual function in 323 consecutive kidney transplant recipients was conducted by mail by means of the validated International Index of Erectile Function (IIEF). All five IIEF domains (IIEF-5), i.e., erectile function, orgasmic function, sexual desire, intercourse satisfaction, and overall satisfaction, were scored for each responder. IIEF-5 scoring that conformed to the National Institutes of Health definition of erectile dysfunction (ED) was computed for all patients sexually active within the past 4 weeks. RESULTS: Two hundred and seventy-one patients replied. Compared to the controls used for IIEF psychometric validation, kidney transplant recipients gave lower erectile function (P<0.01) and intercourse satisfaction (P<0.05) scores, despite their being younger. ED, according to the IIEF-5 method, was demonstrated in 55.7% of the sexually active patients (n=212). Age, time on dialysis, and iterative transplants were significantly and negatively related to erectile dysfunction. CONCLUSION: IIEF proved to be a valuable means of unveiling highly prevalent erectile dysfunction in male kidney transplant recipients. The negative impact of the time on dialysis was emphasized in the results.

Adult↗

Development of an emergency department work score to predict ambulance diversion.

OBJECTIVES: The authors sought to develop and validate an emergency department (ED) work score that could be used in real time to quantify crowding and staff workload in an ED. This work score could be used by public health officials to direct ambulance traffic based on an objective measure of ED status and to track ED conditions over time. In addition, the authors sought to determine which portion of ED care was most responsible for crowding. METHODS: The setting was a tertiary teaching hospital with an emergency medicine residency. A number of ED parameters were measured throughout 2003 and then matched to times that an ED was on diversion status. Odd months of the year were used to develop the standard and even months to validate the standard. A marginal logistic regression analysis was used to develop the standard. The decision to divert ambulances was used as the criterion for ED crowding. RESULTS: The logistic regression demonstrated excellent correlation between the work score and diversion status. At the point of maximum inflection of the receiver operating characteristic curve, the work score predicted diversion status with 86% sensitivity and 80% specificity. CONCLUSIONS: An ED work score was successfully developed and internally validated. External validation should be performed before widespread use.

Ambulances↗

[Validation of a questionnaire for the diagnosis of urinary incontinence].

OBJECTIVE: To validate a questionnaire applied in the Primary Care (PC) clinic which enables urinary incontinence (UI) and its different types to be diagnosed. DESIGN: A descriptive crossover study. SETTING: A Urodynamics hospital out-patients clinic. PARTICIPANTS: Patients referred from PC to be tested for UI by Urodynamics. INTERVENTION: A self-filled questionnaire prior to the Urodynamics test to give a rough idea of the type of UI. Analysis of patients' characteristics and the internal validity of the questionnaire by comparing it with the Urodynamics test. MEASUREMENTS AND MAIN RESULTS: The sample was 59 men and 432 women. For the rough diagnosis of UI caused by straining in women, a five-question survey had a positive predictive value (PPV) of 77.2% for four affirmative replies, which went up to 83% if maximum exterior flow was included. For the UI group arising from anticholinergic treatment, a four-question survey had low PPV (57.6%), but this figure went up to 85.7% with a flowmeter. CONCLUSIONS: A questionnaire to study the type of UI, along with further tests, approached the aetiological diagnosis of incontinence in women. A Urodynamics study, however, always needs to be performed on men.

Adult↗

Informed prognosis [corrected] after abdominal aortic aneurysm repair using predictive modeling techniques [corrected].

OBJECTIVE: To identify the best method for the prediction of postoperative mortality in individual abdominal aortic aneurysm surgery (AAA) patients by comparing statistical modelling with artificial neural networks' (ANN) and clinicians' estimates. METHODS: An observational multicenter study was conducted of prospectively collected postoperative Acute Physiology and Chronic Health Evaluation II data for a 9-year period from 24 intensive care units (ICU) in the Thames region of the United Kingdom. The study cohort consisted of 1205 elective and 546 emergency AAA patients. Four independent physiologic variables-age, acute physiology score, emergency operation, and chronic health evaluation-were used to develop multiple regression and ANN models to predict in-hospital mortality. The models were developed on 75% of the patient population and their validity tested on the remaining 25%. The results from these two models were compared with the observed outcome and clinicians' estimates by using measures of calibration, discrimination, and subgroup analysis. RESULTS: Observed in-hospital mortality for elective surgery was 9.3% (95% confidence interval [CI], 7.7% to 11.1%) and for emergency surgery, 46.7% (95% CI, 42.5 to 51.0%). The ANN and the statistical models were both more accurate than the clinicians' predictions. Only the statistical model was internally valid, however, when applied to the validation set of observations, as evidenced by calibration (Hosmer-Lemeshow C statistic, 14.97; P = .060), discrimination properties (area under receiver operating characteristic curve, 0.869; 95% CI, 0.824 to 0.913), and subgroup analysis. CONCLUSIONS: The prediction of in-hospital mortality in AAA patients by multiple regression is more accurate than clinicians' estimates or ANN modelling. Clinicians can use this statistical model as an objective adjunct to generate informed prognosis.

Aortic Aneurysm, Abdominal↗