PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Internal validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Psychometric assessment of the Multidimensional Quality of Life Questionnaire for Persons with HIV/AIDS (MQOL-HIV) in a sample of HIV-infected women.

Since the late 1980s, several HIV-specific quality of life instruments have been developed; however, little testing has been done in terms of their validity and reliability for HIV-infected women. The purpose of this study was to test the content validity, concurrent validity, internal consistency, and test-retest reliability of the Multidimensional Quality of Life Questionnaire for Persons with HIV/AIDS (MQOL-HIV) in a sample of 85 HIV-infected women. The MQOL-HIV is a 40-item scale comprised of 10 dimensions. Most of the items and all of the domains were determined content valid but revision of some of the items and domains is recommended. Concurrent validity was measured between the MQOL-HIV and the MOS-HIV and ranged from 0.51-0.81 between similar domains. Of the 10 domains and the entire instrument, 7 had a Cronbach's alpha over 0.70 (range 0.43-0.92). Eight domains and the entire instrument achieved test-retest correlation coefficients over 0.70 (range 0.60-0.96). Although some revision may make the scale more content-valid for HIV-infected women, given due care in the interpretation of results, the MQOL-HIV can be used with female populations in its current form.

Adult↗

Development and external validation of an explainable machine learning model for predicting chronic kidney disease progression in the Korean population.

BACKGROUND: Current risk stratification models, such as the Kidney Failure Risk Equation (KFRE), exhibit variable performance across ethnic groups and fail to capture dynamic clinical trajectories. This study aimed to develop and validate a Korean-specific machine learning (ML) model for predicting chronic kidney disease (CKD) progression using an ensemble approach. METHODS: We used electronic health records from Seoul National University Hospital for model development (n = 28,209) and the Korean Genome and Epidemiology Study (KoGES) CKD cohort for external validation (n = 3,960). The primary outcome was a composite of ≥40% decline in estimated glomerular filtration rate (eGFR) or progression to end-stage renal disease within 2 years. A soft-voting ensemble of four ML algorithms (XGBoost, LightGBM, CatBoost, and Random Forest) was developed. RESULTS: The ensemble model demonstrated robust discrimination in internal validation (area under the receiver operating characteristic curve [AUROC], 0.939; 95% confidence interval [CI], 0.934-0.944), significantly exceeding the KFRE (AUROC, 0.879-0.884). External validation in the KoGES cohort showed comparable discrimination (AUROC, 0.859; 95% CI, 0.798-0.914) versus KFRE (four-variable AUROC, 0.882; 95% CI, 0.818-0.935). Shapley Additive exPlanations (SHAP) analysis identified baseline eGFR, serum creatinine, eGFR slope, albumin, and hemoglobin as key prognostic features, supporting a complementary framework using KFRE for community screening and the ML model for hospital-based risk stratification. CONCLUSION: The ensemble ML model accurately predicts short-term CKD progression in Korean patients. By incorporating longitudinal features and ensemble learning, it provides a precise alternative to Western-derived equations, particularly in tertiary care settings.

Chronic kidney failure↗

Duplex imaging immediately prior to carotid endarterectomy.

INTRODUCTION: In this centre, angiography is used only in selected cases, whilst duplex ultrasound (DU) is the main imaging method prior to carotid endarterectomy (CEA). DU has no associated morbidity and so can be repeated immediately before surgery to detect changes in the carotid plaque or degree of stenosis. PATIENTS AND METHODS: We retrospectively examined our Vascular Surgery Audit database for the last 500 patients admitted for CEA. In each case, the DU scan was repeated immediately before surgery. RESULTS: From 500 admissions, repeat DU immediately prior to surgery detected 8 (1.6%) situations where CEA would no longer have been an appropriate intervention. In four cases, the degree of stenosis was found to be less than 70% on the repeat scan - in three cases the internal carotid artery (ICA) had occluded or sub-occluded and in one case there was a dissection of the ICA plaque. CONCLUSIONS: DU can be repeated, with no associated morbidity, immediately prior to surgery. Such a practice changes management decisions in 1.6% of admissions for CEA, allowing surgery unjustified by current evidence to be avoided. This policy also serves several other important purposes: it is a method of internal validation, provides a means of improving training of vascular technologists and of achieving quality assurance in DU techniques.

Carotid Artery, Internal↗

Variations in lung cancer risk among smokers.

BACKGROUND: Although there is no proven benefit associated with screening for lung cancer, screening programs are attracting many individuals who perceive themselves to be at high risk due to smoking. We sought to determine whether the risk of lung cancer varies predictably among smokers. METHODS: We used data on 18 172 subjects enrolled in the Carotene and Retinol Efficacy Trial (CARET)-a large, randomized trial of lung cancer prevention-to derive a lung cancer risk prediction model. Model inputs included the subject's age, sex, asbestos exposure history, and smoking history. We assessed the model's calibration by comparing predicted and observed rates of lung cancer across risk deciles and validated it by assessing the extent to which a model estimated on data from five CARET study sites could predict events in the sixth study site. We then applied the model to evaluate the risk of lung cancer among smokers enrolled in a study of lung cancer screening with computed tomography (CT). RESULTS: The model was internally valid and well calibrated. Ten-year lung cancer risk varied greatly among participants in the CT study, from 15% for a 68-year-old man who has smoked two packs per day for 50 years and continues to smoke, to 0.8% for a 51-year-old woman who smoked one pack per day for 28 years before quitting 9 years earlier. Even among the subset of CT study participants who would be eligible for a clinical trial of cancer prevention, risk varied greatly. CONCLUSIONS: The risk of lung cancer varies widely among smokers. Accurate risk prediction may help individuals who are contemplating voluntary screening to balance the potential benefits and risks. Risk prediction may also be useful for researchers designing clinical trials of lung cancer prevention.

Aged↗

Analysis of case-only studies accounting for genotyping error.

The case-only design provides one approach to assess possible interactions between genetic and environmental factors. It has been shown that if these factors are conditionally independent, then a case-only analysis is not only valid but also very efficient. However, a drawback of the case-only approach is that its conclusions may be biased by genotyping errors. In this paper, our main aim is to propose a method for analysis of case-only studies when these errors occur. We show that the bias can be adjusted through the use of internal validation data, which are obtained by genotyping some sampled individuals twice. Our analysis is based on a simple and yet highly efficient conditional likelihood approach. Simulation studies considered in this paper confirm that the new method has acceptable performance under genotyping errors.

Bias↗

Mortality prediction using SAPS II: an update for French intensive care units.

INTRODUCTION: The standardized mortality ratio (SMR) is commonly used for benchmarking intensive care units (ICUs). Available mortality prediction models are outdated and must be adapted to current populations of interest. The objective of this study was to improve the Simplified Acute Physiology Score (SAPS) II for mortality prediction in ICUs, thereby improving SMR estimates. METHOD: A retrospective data base study was conducted in patients hospitalized in 106 French ICUs between 1 January 1998 and 31 December 1999. A total of 77,490 evaluable admissions were split into a training set and a validation set. Calibration and discrimination were determined for the original SAPS II, a customized SAPS II and an expanded SAPS II developed in the training set by adding six admission variables: age, sex, length of pre-ICU hospital stay, patient location before ICU, clinical category and whether drug overdose was present. The training set was used for internal validation and the validation set for external validation. RESULTS: With the original SAPS II calibration was poor, with marked underestimation of observed mortality, whereas discrimination was good (area under the receiver operating characteristic curve 0.858). Customization improved calibration but had poor uniformity of fit; discrimination was unchanged. The expanded SAPS II exhibited good calibration, good uniformity of fit and better discrimination (area under the receiver operating characteristic curve 0.879). The SMR in the validation set was 1.007 (confidence interval 0.985-1.028). Some ICUs had better and others worse performance with the expanded SAPS II than with the customized SAPS II. CONCLUSION: The original SAPS II model did not perform sufficiently well to be useful for benchmarking in France. Customization improved the statistical qualities of the model but gave poor uniformity of fit. Adding simple variables to create an expanded SAPS II model led to better calibration, discrimination and uniformity of fit, producing a tool suitable for benchmarking.

Adult↗

The Major Depression Rating Scale (MDS). Inter-rater reliability and validity across different settings in randomized moclobemide trials. Danish University Antidepressant Group.

The Major Depression Rating Scale (MDS) has been derived from the Hamilton Depression Scale and the Melancholia Scale. The MDS contains the nine DSM-IV items for major depression which all have anchoring scores from 0 to 4; hence, the theoretical score range is up to 36. The Major Depression Rating Scale has in this study been psychometrically analysed in randomized moclobemide trials. The results showed that the MDS had higher internal validity than the Hamilton Depression Scale. Thus, the homogeneity of the items was higher; factor analysis identified only one general depression factor (after 4 weeks of treatment explaining more than 50% of the variance). The inter-rater reliability of the two scales was of the same high level. The ability to measure changes (external validity) was tested in randomized clinical trials with moclobemide versus tricyclics (clomipramine and notriptyline) performed in Denmark in the psychiatric setting as well as in the general practice. The results showed that in the psychiatric setting tricyclics were superior to moclobemide with effect sizes ranging between 0.43 and 0.53. The highest effect size was obtained with the Melancholia Scale and the Major Depression Rating Scale, while the Hamilton Depression Scale was below 0.50. In the general practice setting no difference was found between moclobemide and clomipramine. In conclusion, the Major Depression Rating Scale has been found to have a more homogeneous factor structure than the Hamilton Depression Scale, but still with the same level of reliability and external validity. However, studies are needed to standardize the scale, especially in the general practice setting.

Antidepressive Agents↗

Prevalence, incidence, signs and treatment of clinical listeriosis in dairy cattle in England.

The prevalence, incidence and clinical signs of listeriosis in dairy cattle in England were investigated by means of a postal questionnaire survey of 1500 dairy farmers. The response rate was 64.1 per cent. Overall the farm prevalence of listeriosis was 11.7 per cent, 9.3 per cent for milking cows, 5.0 per cent for replacement heifers and 1.4 per cent for dairy calves. The within-herd incidence rate per thousand animal-years was 51.4 for all cases, 39.7 for milking cows, 86.6 for replacement heifers and 73.7 for dairy calves. Most cases of clinical listeriosis were reported between December and May, and the most common signs were silage eye, followed by nervous signs. The results of the questionnaire were validated internally by re-estimating the farm prevalence by including only those cases diagnosed by a veterinarian or veterinary investigation centre; the prevalence did not change significantly. The proportion of cases which were culled or died of encephalitic listeriosis was compared with the proportion diagnosed during statutory BSE reporting. The fact that the two proportions were similar provided external validation for the results of the questionnaire.

Animals↗

Carotid duplex imaging: variation and validation.

BACKGROUND: Duplex imaging is increasingly used as the only investigation before carotid endarterectomy, but many different criteria exist in the literature for the detection of a severe (70-99 per cent) carotid stenosis. This study aimed to investigate current practice in carotid duplex imaging in Great Britain and Ireland. METHODS: A postal questionnaire was sent to 86 vascular surgical units. RESULTS: The median number of scans performed per year was 450 (range 60-4500). Thirty-six per cent of units who responded used peak systolic : end diastolic velocity ratio to calculate carotid stenosis. Overall, nine different major duplex criteria were used to grade carotid stenosis in 14 different systems of percentage bands. Only 51 per cent of units verified their duplex criteria against angiography. Eighteen per cent of units used two or more different types of duplex scanner and applied the same diagnostic criteria to each machine. CONCLUSION: A wide variation in diagnostic duplex criteria and methods of grading stenosis exists among vascular units. Internal validation is not performed routinely. Standardization of duplex criteria would ensure greater consistency, but would not replace the need for validation of results within each unit.

Carotid Artery, Internal↗

A short version of the Self Description Questionnaire II: operationalizing criteria for short-form evaluation with new applications of confirmatory factor analyses.

Four studies evaluate the new Self Description Questionnaire II short-form (SDQII-S) that measures 11 dimensions of adolescent self-concept based on responses to 51 of the original 102 SDQII items and demonstrate new statistical strategies to operationalize guidelines for short-form evaluation proposed by G. T. Smith, D. M. McCarthy, and K. G. Anderson (2000). Multiple-group confirmatory factor analyses revealed that the factor structure based on responses to 51 items by a new cross-validation group (n=9,134) was invariant with the factor structures based on responses to the same 51 items and to all 102 items by the original normative archive group (n = 9,187). Reliabilities for the 11 SDQII-S factors were nearly the same and consistently high (.80 to .89) for both groups. Multitrait-multimethod analyses support the internal validity of responses over time. Gender and age effects on the 11 SDQII-S factors were invariant across the archive and cross-validation groups.

Adolescent↗

Clinical trials of primary care treatments for major depression: issues in design, recruitment and treatment.

The objective of this article is to consider whether randomized clinical trials (RCTs) are able to determine the validity of transferring treatments for major depression from the psychiatric to the primary care sector. This clinical issue is of growing concern in the United States since both governmental and professional bodies are establishing guidelines for the treatment of medical patients with the affective disorder. The article's method involves analysis of how the competing aims of rigorous scientific methodology (internal validity) and generalization of study findings (external validity) are best balanced within the RCT. Experiences in recruiting medical patients with major depression and providing pharmacologic, psychotherapeutic, and usual care interventions compatible with the sociotechnical characteristics of ambulatory medical centers are described to illustrate the complexities of investigating transferability of treatments for major depression with RCT methodology.

Antidepressive Agents↗

Quality of life in patients with oesophageal cancer.

There is a growing interest in assessing quality of life in patients with oesophageal cancer because it provides detailed information of the patients' perception of the benefits or harms of treatment. Yet few studies have prospectively measured quality of life using validated appropriate instruments. There are now several questionnaires for patients with cancer, although these are not sufficiently sensitive to small but clinically important changes in quality of life. It is therefore recommended that a disease-specific module is used in conjunction with generic measures. The European Organisation into Research and Treatment of Cancer (EORTC) QLQ-OES24 is currently completing an international validation study. It is used with the EORTC QLQ-C30 core instrument and is designed for patients undergoing potentially curative treatment or palliation of malignant dysphagia. Studies that have assessed quality of life after oesophagectomy have generally found that survivors do regain their former health. Little is known about the effect of neoadjuvant chemoradiation on patients' quality of life. Following endoscopic palliation of dysphagia, quality of life can be maintained and improvement of swallowing is seen. A validated appropriate assessment of quality of life should be included in future palliative trials and in studies of new treatments which may marginally influence survival but cause significant side effects.

Deglutition Disorders↗

Development of a model for case-mix adjustment of pressure ulcer prevalence rates.

BACKGROUND: Acute care hospitals participating in the Dutch national pressure ulcer prevalence survey use the results of this survey to compare their outcomes and assess their quality of care regarding pressure ulcer prevention. The development of a model for case-mix adjustment is essential for the use of these prevalence rates as an outcome measure. OBJECTIVE: The development of a valid model for case-mix adjustment to compare the prevalence rates in the acute care hospitals that participated in the 1998 Dutch pressure ulcer prevalence survey, for the purpose of performance comparisons among the hospitals. DESIGN: Cross-sectional design. SUBJECTS: Subjects were patients residing in the 43 acute care hospitals that participated in the national pressure ulcer prevalence survey on May 26, 1998. MEASURES: The study examined the validity of a model for case-mix adjustment of pressure ulcer prevalence rates and compared hospitals to evaluate the impact of adjusted prevalence rates on their performance. RESULTS: A logistic model was developed for case-mix adjustment, using age, malnutrition, incontinence, activity, mobility, sensory perception, friction and shear, and ward specialty. This model was found to have content, construct, and internal validity. Case-mix adjustment influenced the hospitals' performance. CONCLUSION: The data of the national pressure ulcer prevalence survey can be used to develop a valid model for case-mix adjustment. Conclusions about the quality of care were influenced by the use of case-mix adjusted outcomes as a measure of this quality.

Adolescent↗

Essential requirements for practice guidelines at national and local levels.

We explored whether local practice guidelines (PGs) on stroke management had undergone a process of local adaptation in relation to the appropriate and feasible configuration of stroke units (SUs). We critically appraised 7 PGs developed by 6 Italian local healthcare units, using explicit criteria to evaluate internal validity and their adequacy relative to local implementation issues. All PGs were developed by multidisciplinary working groups. In 4 of 6 PGs recommending SUs for stroke care, methodology for evidence retrieval was poor. Although organisational aspects were addressed in 4 of 6 PGs, details on how a SU should be organised were not provided in any of the examined PGs. Despite availability of national and international stroke PGs, at local level guidelines developers seem to spend time in "reinventing the wheel" rather than concentrating on what matters for local implementation. Besides being inefficient, this seems to lead to methodologically poor products inappropriate for what should be done to assure that interventions that work are packaged in a way that is compatible with their uptake into the ongoing services activities.

Hospital Units↗

Common sense and figures: the rhetoric of validity in medicine (Bradford Hill Memorial Lecture 1999).

Austin Bradford Hill was once a friend to The Lancet, but, as occasionally happens, friends fall out. The great legacy of his association with the journal, however, was Principles of Medical Statistics. As each edition was succeeded by another--the first in 1937, the last in 1991--he seemed to shift his view about the influence of statistical method on clinical practice from one of assured certainty to one of modest advantage. That change paralleled a move away from an emphasis on the importance of internal validity in the randomized trial to one of understanding the inescapably practical significance of generalizability. Writers on medical research have explored notions of external validity in various ways. One view, for example, is to seek a close correlation between the participants in a clinical trial and patients seen in practice. The argument goes that such a correspondence has to be made before any decision can be taken about whether to apply the result of that trial to the clinical setting. Another view, first worked out by the American logician Charles Sanders Peirce, is that one must simply rely on the informed guess, based on a reasonable estimate of the limits of extrapolation. The tensions between and implications of these two different approaches are worked through using the example of coronary stents. A solution is, perhaps, to write explicit rules of interpretation that provide a framework for judging the strength of a claim to applicability. Five questions are posed, which try to lay a foundation for such a framework.

Clinical Medicine↗

Cardiac disease outcomes in clinical trials.

Randomized controlled trials (RCTs) are the gold standard for determining causality in medicine. To be considered valid, the efficacy/effectiveness of a new drug must be tested and proved through this type of scientific study. This type of trial will disclose the drug's risk profile as well as the treatment effect magnitude. The proper design, development, and analysis and presentation of results from an RCT is based on a group of well-defined methodological rules, compliance with which assures the trial's internal validity, the relative and absolute importance of the results and its applicability to populations of patients different from those included in the study sample (external validity). Among the structural and methodological components of a clinical trial--randomization, allocation concealment, confounding, similarity of study groups, measures of efficacy, statistical analysis, etc.--disease markers (endpoints, outcomes) are especially important. In the end, what an RCT is good for is to detect changes in the disease process with therapy (or preventive measures), and these changes are defined beforehand based on specific measurements--disease markers. In this paper we will present general principles for the definition of disease markers, their problems and practical use. Caution should be exercised, as this is an area of clinical epidemiology that is somewhat complex, controversial and ill-defined.

Databases, Factual↗

International nursing research in social support: theoretical and methodological issues.

Social support has been widely studied within cultures as a variable that is protective of health and mental health, either directly or as a buffer against life stress. When research is conducted across cultures, several conceptual and methodological issues emerge. The purpose of this paper is to discuss the conceptual basis for assuming that social support is a universal phenomenon, to suggest areas in which manifestations of social support may be culture-specific, and to present methodological issues that need to be addressed in conducting valid international research on social support.

Cross-Cultural Comparison↗

Can we use contingent valuation to assess the demand for childhood immunisation in developing countries?: a systematic review of the literature.

Childhood immunisation is one of the most cost-effective public health interventions, yet its population coverage in low- and middle-income countries is severely limited by the fiscal constraints that health services face. A recent proposal suggested that commitments to purchase vaccines and make them available to developing countries for modest co-payments could solve the problem. However, this is dependent on communities being willing and able to share the cost in this way, which is difficult to assess. One possible method to assess this demand is contingent valuation (CV). This article evaluates the usefulness of using CV in this way, by reviewing applications of CV in developing countries against current 'standards' for CV of immunisation in the literature. A structured review was adopted with reference to the standard frameworks for methodological evaluation. A set of five criteria were developed for evaluating an 'acceptable' CV study: (i) response rate; (ii) association between willingness to pay (WTP) and socioeconomic status (SES); (iii) sensitivity of WTP to benefit scale/scope; (iv) predictive validity; and (v) reliability in elicitation formats. Two strands of literature search were conducted using electronic databases (MEDLINE, EMBASE, HEALTHSTAR and Econlit) from 1966 to 2003, one for CV studies of immunisation and one for CV studies in developing countries. Twelve CV studies of vaccination and 13 CV studies undertaken within developing countries were identified and reviewed. The quality of existing CV studies conducted in developing countries exceeded the benchmark standard set by studies of immunisation in the developed world in four of the five criteria. WTP estimates appeared both internally valid (i.e. associations with SES) and externally valid (i.e. predictive validity), reliability in developing countries was no less than that of the benchmark level in the existing literature, and the high response rates suggested that CV can be administered to a rural, and perhaps less literate, population. Only sensitivity to scale/scope was not well demonstrated. Our assessment indicated that the CV technique offers a promising tool to estimate the demand for childhood immunisation in low- and middle-income countries. International agencies are therefore encouraged to devote resources to such an application when designing their support to the immunisation programmes.

Benchmarking↗