PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “external validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Clinical trials and external validity.

Although a number of clinical trials are available estimating the benefits of lipid-lowering therapies that include economic end-points, development of modeling methodologies are essential to extend results of those trials over time and to other populations. We reviewed the key issues to be considered when extending trial data to real-world situations. The availability of recent randomized controlled trials of 3-hydroxy-3-methylglutaryl coenzyme A (HMG-CoA) reductase inhibitors in primary and secondary prevention has demonstrated the limitations of earlier modeling efforts to project benefits of lipid modification. The importance of risk stratification is demonstrated, particularly the importance of both LDL and HDL cholesterol either together or as a ratio measure. The selection of modeling methodology to extend benefits of a treatment beyond the end of a trial and over a lifetime is discussed. The relationship between benefits from lipid reduction and risk difference is described, demonstrating that for individuals with established coronary heart disease (CHD) and those older than age 58, benefits from lipid reduction are greater than those predicted from baseline lipid-related risk differences alone. The implications of these data for primary prevention of CHD in the elderly are discussed.

Journal Article↗

External validation of EPIWIN biodegradation models.

The BIOWIN biodegradation models were evaluated for their suitability for regulatory purposes. BIOWIN includes the linear and non-linear BIODEG and MITI models for estimating the probability of rapid aerobic biodegradation and an expert survey model for primary and ultimate biodegradation estimation. Experimental biodegradation data for 110 newly notified substances were compared with the estimations of the different models. The models were applied separately and in combinations to determine which model(s) showed the best performance. The results of this study were compared with the results of other validation studies and other biodegradation models. The BIOWIN models predict not-readily biodegradable substances with high accuracy in contrast to ready biodegradability. In view of the high environmental concern of persistent chemicals and in view of the large number of not-readily biodegradable chemicals compared to the readily ones, a model is preferred that gives a minimum of false positives without a corresponding high percentage false negatives. A combination of the BIOWIN models (BIOWIN2 or BIOWIN6) showed the highest predictive value for not-readily biodegradability. However, the highest score for overall predictivity with lowest percentage false predictions was achieved by applying BIOWIN3 (pass level 2.75) and BIOWIN6.

Biodegradation, Environmental↗

External validation of prognostic models for ongoing pregnancy after in-vitro fertilization.

This study aimed to validate prognostic models for predicting ongoing pregnancy after the first and second in-vitro fertilization cycles. Models were developed using data from the University Hospital, Nijmegen, 1991-1994 and tested using more recent data from the same centre and data from two other centres. Although the variables included in the models seemed plausible, the predictions of the models were unsatisfactory. The models did not discriminate between women who had achieved pregnancy and women who did not achieve pregnancy; neither could they indicate which women had a (very) low probability of ongoing pregnancy. Taking into account the success rate of a specific clinic or the success rate during a specific period did not show any advantage. The predictions were even inaccurate in the same hospital during another period. It is obvious that these prognostic models should not be used. This study shows the importance of validating prognostic models before their implementation in clinical practice.

Adult↗

Pregnancy is predictable: a large-scale prospective external validation of the prediction of spontaneous pregnancy in subfertile couples.

BACKGROUND: Prediction models for spontaneous pregnancy may be useful tools to select subfertile couples that have good fertility prospects and should therefore be counselled for expectant management. We assessed the accuracy of a recently published prediction model for spontaneous pregnancy in a large prospective validation study. METHODS: In 38 centres, we studied a consecutive cohort of subfertile couples, referred for an infertility work-up. Patients had a regular menstrual cycle, patent tubes and a total motile sperm count (TMC) >3 x 10(6). After the infertility work-up had been completed, we used a prediction model to calculate the chance of a spontaneous ongoing pregnancy (www.freya.nl/probability.php). The primary end-point was time until the occurrence of a spontaneous ongoing pregnancy within 1 year. The performance of the pregnancy prediction model was assessed with calibration, which is the comparison of predicted and observed ongoing pregnancy rates for groups of patients and discrimination. RESULTS: We included 3021 couples of whom 543 (18%) had a spontaneous ongoing pregnancy, 57 (2%) a non-successful pregnancy, 1316 (44%) started treatment, 825 (27%) neither started treatment nor became pregnant and 280 (9%) were lost to follow-up. Calibration of the prediction model was almost perfect. In the 977 couples (32%) with a calculated probability between 30 and 40%, the observed cumulative pregnancy rate at 12 months was 30%, and in 611 couples (20%) with a probability of >or=40%, this was 46%. The discriminative capacity was similar to the one in which the model was developed (c-statistic 0.59). CONCLUSIONS: As the chance of a spontaneous ongoing pregnancy among subfertile couples can be accurately calculated, this prediction model can be used as an essential tool for clinical decision-making and in counselling patients. The use of the prediction model may help to prevent unnecessary treatment.

Adult↗

External validation of severity scoring systems for acute renal failure using a multinational database.

OBJECTIVE: Several different severity scoring systems specific to acute renal failure have been proposed. However, most validation studies of these scoring systems were conducted in a single center or in a small number of centers, often the same ones used for their development. Therefore, it is not known whether such severity scoring systems may be widely applied. DESIGN: Prospective clinical investigation. SETTING: Intensive care units. PATIENTS: One thousand seven hundred and forty-two intensive care unit patients with acute renal failure who were either treated with renal replacement therapy or fulfilled predefined criteria. INTERVENTIONS: Demographic and clinical information and outcomes were measured. MEASUREMENTS AND MAIN RESULTS: Scores for four acute renal failure-specific scoring systems and two general scoring systems (Simplified Acute Physiology Score II and Sequential Organ Failure Assessment) were calculated, and their discrimination and calibration were tested with receiver operating characteristic curves and Hosmer-Lemeshow goodness-of fit-tests. For the receiver operating characteristic curves, blood lactate levels were also used as a reference. All scores had an area under the receiver operating characteristic curve <0.7 (Mehta 0.670, Liano 0.698, Chertow 0.610, Paganini 0.643, Simplified Acute Physiology Score II 0.645, Sequential Organ Failure Assessment 0.675, lactate 0.639). For scores that can calculate predicted mortality, the Hosmer-Lemeshow goodness-of-fit test showed poor calibration. CONCLUSIONS: None of the scoring systems tested had a high level of discrimination or calibration to predict mortality for patients with acute renal failure when tested in a broad cohort of patients from multiple countries. A large, multiple-center database might be needed to improve the discrimination and calibration of acute renal failure scoring system.

Acute Kidney Injury↗

A comparison of two primary care trials on tennis elbow: issues of external validity.

OBJECTIVE: To assess clinical heterogeneity across two studies with respect to study population, interventions, and outcome measures, and to evaluate the influence of these sources of heterogeneity on the results of the studies. METHODS: The individual patient data were used from two randomised controlled trials investigating the effectiveness of conservative treatments in patients with tennis elbow in primary care. Patients were allocated at random to treatment with steroid injection, wait and see policy, non-steroidal anti-inflammatory drugs, placebo tablets, or physiotherapy. Outcome measures included severity of the main complaint, inconvenience of the elbow complaints, pain during the day, elbow disability, pain-free grip strength, and global improvement. All outcomes were assessed at 1, 6, and 12 months after randomisation. RESULTS: The two study populations were similar with respect to age, sex, comorbid neck/shoulder complaints, and baseline scores for the severity of pain. However, significant differences were observed for employment status, duration of elbow complaints, dominant side affected, previous history of elbow complaints, and use of analgesics. Local injections differed between the two studies with respect to volume, number, and steroid preparation. However, after 1, 6, and 12 months, the treatment effects of steroid injections were very similar between the study populations. CONCLUSIONS: Despite large differences in study population at baseline, the responses to steroid injections were remarkably similar. Also the responses to other conservative interventions and the placebo treatment were very consistent, suggesting a uniform course of a tennis elbow and a lack of influence of clinical heterogeneity.

Adult↗

Carotid artery stenosis: external validity of the North American Symptomatic Carotid Endarterectomy Trial measurement method.

PURPOSE: To assess the generalizability of the North American Symptomatic Carotid Endarterectomy Trial method for determining the degree of stenosis on angiograms. MATERIALS AND METHODS: Six good-quality, baseline angiograms of carotid arteries that were less than 70% stenosed were reviewed by 14 experienced neuroradiologists at different academic institutions. All reviewers determined the degree of stenosis by calculating the ratio of the diameter of the artery at the point of maximal narrowing to the normal diameter distal to the stenosis (well beyond the carotid artery bulb). The reviewers marked the location of their measurements on the angiogram. Comparisons were performed among the reviewers' results and with the reference measurements. RESULTS: Interobserver agreement was 0.84 (95% confidence interval = 0.65, 0.97). The average interobserver disagreement of +/-7% was comparable with that reported in the literature. The overall bias was 6%, which indicated a tendency of the reviewers to overestimate the degree of stenosis in comparison with the reference determination. CONCLUSION: The North American Symptomatic Carotid Endarterectomy Trial reference measurements can be generalized beyond the bounds of this clinical trial, provided that attention is paid to details of the measurement method.

Angiography↗

Scoring patient management problems: external validation of expert consensus.

This study determined the extent to which medical faculty agreed when rating options in two written patient management problems in diabetes mellitus. Another purpose was to determine whether option weights (used for scoring) based on the consensus ratings of faculty actually predict the choices of well-qualified physicians. Experts showed better-than-chance agreement but with considerable variation from one part of a problem to another (31% to 73%). Nevertheless, consensus ratings were very accurate predictors of the decisions of endocrinology fellows (correlations from .60 to .97). When scoring weights are assigned to options in patient management problems, consensus (average) ratings of experts are likely to demonstrate high concurrent validity for well-qualified clinicans.

Adult↗