PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “External validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

[Relationship between prothrombin time international normalized ratio and thrombo test (%)].

OBJECTIVES: The optimal therapeutic range for laboratory evaluation of oral anticoagulant therapy is now defined by the prothrombin time international normalized ratio (PT-INR). However, the thrombo test (TT), an alternative method to measure intensity of anticoagulation, is also currently used throughout Japan. The relationship between PT-INR and TT (%) has yet to be clarified. This study investigated the relationship between PT-INR and TT (%). METHODS: The PT-INR and TT (%) were simultaneously measured of 505 consecutive samples from patients treated with warfarin in our hospital. Fourteen functions were used for regression analyses: a fractional function (Y = a/X + b), a square root function (Y = aX0.5 + b), a natural logarithmic function (Y = a.lnX + b), a power series function (Y = aXb), a quotient function (Y = abX), and polynomial functions [Y = anXn + an - 1Xn - 1 +......+ a1X1 + b, (1 < or = n < or = 9)]. The results were confirmed by the same methods in 383 samples and 296 samples from another two laboratories. RESULTS: The power series function showed the most significant (p < 0.0001) and highest adjusted R2 (0.858) correlation, with a regression formula of TT (%) = e4.48 (PT-INR)-2.09 in our laboratory. Using the same analyses, the power series function also showed the most significant and highest adjusted R2 in samples from the other two laboratories. CONCLUSIONS: This study showed that a power series function is the most appropriate for expressing the relationship between PT-INR and TT (%) among the 14 functions. The function between PT-INR and TT (%) is mainly derived from the relationship between TT (%) and TT (sec). Both internal validity and external validity confirmed the relationship between PT-INR and TT (%).

Aged↗

A methodological framework for the design of research on the evaluation of residents.

This paper describes a construct validation framework for research on the selection and evaluation of residents. The application of the proposed methodology to surgery residents is described. The need to measure non-cognitive and neuropsychological factors in addition to cognitive knowledge and technical ability is emphasized, and a research strategy that integrates theory formulation, internal validation, and external validation is presented. In this context, residents' competence is viewed as a multivariate construct that requires validation through longitudinal empirical studies and the use of multivariate statistical approaches.

Clinical Competence↗

Scientific challenges in the application of randomized trials.

In recent years, scientific challenges in the application of randomized trials have become more apparent, especially with the extension of such trials to the assessment of nondrug treatments, such as health education, psychotherapy, and health care provision. Six issues (individual v group randomization, blinding and unblinding, the effect of trial participation on outcome, selective subject participation, treatment compliance, and standardized v individualized treatment) are discussed in terms of their impact on internal validity, generalizability (external validity), and clinical relevance. Specific design strategies may be necessary to enhance these methodological and clinical desiderata. Attention to these challenges should lead to improvements in future randomized trials.

Attitude↗

QSAR modeling based on structure-information for properties of interest in human health.

The development of QSAR models based on topological structure description is presented for problems in human health. These models are based on the structure-information approach to quantitative biological modeling and prediction, in contrast to the mechanism-based approach. The structure-information approach is outlined, starting with basic structure information developed from the chemical graph (connection table). Information explicit in the connection table (element identity and skeletal connections) leads to significant (implicit) structure information that is useful for establishing sound models of a wide range of properties of interest in drug design. Valence state definition leads to relationships for valence state electronegativity and atom/group molar volume. Based on these important aspects of molecules, together with skeletal branching patterns, both the electrotopological state (E-state) and molecular connectivity (chi indices) structure descriptors are developed and described. A summary of four QSAR models indicates the wide range of applicability of these structure descriptors and the predictive quality of QSAR models based on them: aqueous solubility (5535 chemically diverse compounds, 938 in external validation), percent oral absorption (%OA, 417 therapeutic drugs, 195 drugs in external validation testing), AMES mutagenicity (2963 compounds including 290 therapeutic drugs, 400 in external validation), fish toxicity (92 substituted phenols, anilines and substituted aromatics). These models are established independent of explicit three-dimensional (3-D) structure information and are directly interpretable in terms of the implicit structure information useful to the drug design process.

Animals↗

Science, ethnicity, and bias: where have we gone wrong?

The quality, quantity, and funding of ethnic minority research have been inadequate. One factor that has contributed to this inadequacy is the practice of scientific psychology. Although principles of psychological science involve internal and external validity, in practice psychology emphasizes internal validity in research studies. Because many psychological principles and measures have not been cross-validated with different populations, those conducting ethnic minority research often have a more difficult time demonstrating rigorous internal validity. Thus, psychology's overemphasis of internal as opposed to external validity has differentially hindered the development of ethnic minority research. To develop stronger research knowledge on ethnic minority groups, it is important that (a) all research studies address external validity issues and explicitly specify the populations to which the findings are applicable; (b) different research approaches, including the use of qualitative and ethnographic methods, be appreciated; and (c) the psychological meaning of ethnicity or race be examined in ethnic comparisons.

Bias↗

Sherlock Holmes and child psychopathology assessment approaches: the case of the false-positive.

OBJECTIVE: To explore the relative value of various methods of assessing childhood psychopathology, the authors compared 4 groups of children: those who met criteria for one or more DSM diagnoses and scored high on parent symptom checklists, those who met psychopathology criteria on either one of these two assessment approaches alone, and those who met no psychopathology assessment criterion. METHOD: Parents of 201 children completed the Child Behavior Checklist (CBCL), after which children and parents were administered the Diagnostic Interview Schedule for Children (version 2.1). Children and parents also completed other survey measures and symptom report inventories. The 4 groups of children were compared against "external validators" to examine the merits of "false-positive" and "false-negative" cases. RESULTS: True-positive cases (those that met DSM criteria and scored high on the CBCL) differed significantly from the true-negative cases on most external validators. "False-positive" and "false-negative" cases had intermediate levels of most risk factors and external validators. "False-positive" cases were not normal per se because they scored significantly above the true-negative group on a number of risk factors and external validators. A similar but less marked pattern was noted for "false-negatives." CONCLUSIONS: Findings call into question whether cases with high symptom checklist scores despite no formal diagnoses should be considered "false-positive." Pending the availability of robust markers for mental illness, researchers and clinicians must resist the tendency to reify diagnostic categories or to engage in arcane debates about the superiority of one assessment approach over another.

Adaptation, Psychological↗

Behavior change intervention research in community settings: how generalizable are the results?

This review examines the extent to which recent behavioral intervention studies conducted in community settings reported on elements of internal and external validity, with an emphasis on whether research has been conducted in representative settings with representative populations. A targeted review was conducted on community-based intervention studies that promoted good nutrition, physical activity or smoking cessation/prevention, and were published in 11 leading health behavior journals between 1996 and 2000. The RE-AIM framework (reach, efficacy, adoption, implementation and maintenance) was used to evaluate the extent to which each paper reported on elements of reach, efficacy/effectiveness, adoption, implementation and maintenance. A total of 27 publications were reviewed. Although most studies (88%) reported participation rates among eligible members of the target audience ('reach'), only 11% of studies reported the participation rate ('adoption') among eligible community-based organizations or settings. Few studies reported if participating individuals or settings were representative of those found in the broader population. Although a majority of studies (59%) reported whether the intervention was delivered ('implementation'), few reported whether individuals maintained behavior change (30%) or whether organizations maintained or institutionalized interventions (0%). To increase the potential to translate community research findings to practice, studies should place a greater emphasis on obtaining and reporting external validity information, such as representativeness. The lack of external validity information limits researchers' and practitioners' ability to judge the generalizability of effects and the comparative utility of interventions. Improved reporting will facilitate implementation of proven and broadly applicable intervention strategies in communities. To make significant progress, all parties, including researchers, reviewers, editors and funders, need to take responsibility for increased emphasis on external validity information and ask what role they can best play to facilitate this process.

Adolescent↗

[Analysis of the scientific evidence of the combination therapy in benign prostatic hyperplasia].

OBJECTIVE: The analysis of the scientific evidence of the combination therapy in Benign Prostatic Hyperplasia (BPH). METHODS: 5 published studies about combination therapy in BPH were analysed following the criteria of the Evidence Based Medicine (EBM). Hypothesis, variables, internal validity, results relevance, and the external validity of every study were analysed. RESULTS: Symptoms changes and maximal flow rate (Qmax) improvement were evaluated in four studies and only one analysed the BPH progression. Inclusion and exclusion criteria were similar in the 5 studies. The main reasons of a poor internal validity in 3 studies were the short follow-up, the missing percentage, the absence of a placebo group and a dosage bias. External validity were decreased in the 5 studies by the exclusion criteria and in 3 of them because high doses of alpha-blocker were given to achieve a therapeutic effect. The study with the highest scientific evidence (MTOPS) is the only that offers confidence intervals and number needed to treat. CONCLUSIONS: The Qmax and the symptoms improvement found in all the studies has a moderate clinical relevance. MTOPS study shows that BPH progression has a low incidence that can be highly reduced by means of combination therapy.

Adrenergic alpha-Antagonists↗

Development and validation of a nomogram predicting the outcome of prostate biopsy based on patient age, digital rectal examination and serum prostate specific antigen.

PURPOSE: We developed and validated a nomogram which predicts presence of prostate cancer (PCa) on needle biopsy. MATERIALS AND METHODS: We used 3 cohorts of men who were evaluated with sextant biopsy of the prostate and whose presenting prostate specific antigen (PSA) was not greater than 50 ng/ml. Data from 4,193 men from Montreal, Canada were used to develop a nomogram based on age, digital rectal examination (DRE) and serum PSA. External validation was performed on 1,762 men from Hamburg, Germany. Data from these men were subsequently used to develop a second nomogram in which percent free PSA (%fPSA) was added as a predictor. External validation was performed using 514 men from Montreal. Both nomograms were based on multivariate logistic regression models. Predictive accuracy was evaluated with areas under the receiver operating characteristic curve and graphically with loess smoothing plots. RESULTS: PCa was detected in 1,477 (35.2%) men from Montreal, 739 (41.9%) men from Hamburg and 189 (36.8%) men from Montreal. In all models all predictors were significant at 0.05. Using age, DRE and PSA external validation AUC was 0.69. Using age, DRE, PSA and %fPSA external validation AUC was 0.77. CONCLUSIONS: A nomogram based on age, DRE, PSA and %fPSA can highly accurately predict the outcome of prostate biopsy in men at risk for PCa.

Adolescent↗

Could the preoperative urethral curve be used to predict immediate urinary continence following Retzius-sparing robot-assisted radical prostatectomy? A retrospective multi-center study.

PURPOSE: Immediate urinary continence (UC) recovery following Retzius-sparing robot-assisted radical prostatectomy (RS-RARP) remains highly variable, highlighting the need for reliable preoperative prediction. We aimed to develop and validate models to identify patients likely to achieve immediate UC recovery following RS-RARP. MATERIALS AND METHODS: A total of 580 prostate cancer patients who underwent RS-RARP from four medical centers were assigned to a training set (n=348), an internal validation set (n=103) and an external validation set (n=129). Independent predictors were identified through univariate analysis and LASSO regression. A nomogram was constructed using multivariate logistic regression. Its performance was evaluated with receiver operating characteristic (ROC) curve, calibration curves, and decision curve analysis. RESULTS: Immediate UC recovery was observed in 84.5% (294/348) of patients in the training cohort, 80.6% (83/103) in the internal validation cohort, and 81.4% (105/129) in the external validation cohort, respectively. Multivariate analysis identified membranous urethral length (MUL) (OR=1.23, P=0.029) and urethral curvature (OR=2.84, P<0.001) as independent predictors, while prostate volume (PV) (OR=0.84, P <0.001) as a protective factor. The nomogram integrating MUL, PV, and urethral curvature demonstrated superior predictive accuracy, with an AUC of 0.87 (95% CI, 0.83-0.91) in the training cohort. The bootstrap-corrected calibration slope was 0.96, and the Brier score was 0.08.&#xa0;Calibration curves and decision curve analysis confirmed the predictive accuracy and clinical utility of the nomogram. CONCLUSIONS: Our study introduces a novel quantitative method for assessing urethral curvature. The mpMRI-based model, integrating urethral curvature and prostate spatial configuration, offers enhanced predictive accuracy for postoperative immediate UC recovery.

Humans↗

Maxillary sinusitis in adults: an evaluation of placebo-controlled double-blind trials.

BACKGROUND: In general practice, acute sinusitis is frequently diagnosed and treated with antibiotics. OBJECTIVE: This study aimed to determine the evidence for the effectiveness of antibiotic treatment in acute maxillary sinusitis in adults by assessing the methodological quality of placebo-controlled double-blind randomized trials. METHOD: An evaluation by four raters through a 35-item scoring-scale for internal and external validity of all placebo-controlled double-blind randomized trials on acute sinusitis found between January 1966 and July 1996. RESULTS: Eighty-five trials were excluded because they were not placebo-controlled, double-blind, randomized, or were carried out in patients with chronic sinusitis or in children. The three remaining trials were performed in different populations (one in general practice) between 1973 and 1978. Only one study claimed superiority of antibiotic treatment. Different inclusion criteria and major outcome measures were used by the authors. The reliability of major outcome events was reported poorly or not at all and in two studies outcome measures were clinically inappropriate. The studies scored 30-62% of the maximum attainable score for internal validity and 10-20% for external validity. CONCLUSION: The effectiveness of antibiotic treatment in acute maxillary sinusitis in a general practice population is not based sufficiently on evidence.

Acute Disease↗

Diagnostic performance of machine learning models versus established risk stratification for intracranial aneurysm rupture: a systematic review and bivariate meta-analysis.

BACKGROUND: Machine learning (ML) models have been proposed to improve the discrimination of intracranial aneurysm rupture status beyond established clinical risk stratification tools. However, reported performance is heterogeneous and the relative contribution of model architecture and feature dominance remains unclear. METHODS: We performed a Preferred Reporting Items for Systematic Reviews and Meta-Analyses-diagnostic test accuracy systematic review and diagnostic meta-analysis of studies evaluating ML models for intracranial aneurysm rupture discrimination. PubMed, Embase and CENTRAL were searched to February 2026. Sensitivity and specificity were pooled using a bivariate random-effects model, with summary receiver operating characteristic curves generated across training, internal testing and external validation datasets. Models were compared with regression-based approaches and Population, Hypertension, Age, Size of aneurysm, Earlier subarachnoid haemorrhage, Site of aneurysm (PHASES) scores. Subgroup and meta-regression analyses explored associations between algorithm family and feature domain. RESULTS: Sixty-two retrospective cohorts (29&#x2009;709 patients 209 models) met the inclusion criteria. In training datasets, pooled sensitivity and specificity for ML were 0.81 (95% CI 0.75 to 0.85)&#x2009;and 0.83 (0.80-0.86), with an area under the curve (AUC) of 0.878, exceeding PHASES (AUC 0.667). In testing datasets, ML retained higher discrimination (AUC 0.837) than regression models (0.806) and PHASES (0.646). In external validation, sensitivity was preserved (0.82), but specificity declined (0.66). Deep learning demonstrated the highest AUCs (training and testing). Incorporation of haemodynamic or radiomic features improved pooled discrimination relative to morphology alone. Evidence of small-study effects and mostly unclear Prediction Model Risk Of Bias Assessment Tool ratings were observed. CONCLUSIONS: ML approaches demonstrate higher pooled discrimination for aneurysm rupture status than conventional risk scores in retrospective datasets, but reduced external validation specificity and heterogeneity limit confidence for clinical translation. Prospective, externally validated, calibrated models are required before integration into routine cerebrovascular risk stratification.

Humans↗

Prediction of incident heart failure in established atherosclerotic cardiovascular disease: the SMART2-HF model.

BACKGROUND AND AIMS: Patients with established atherosclerotic cardiovascular disease (ASCVD) are at high risk of developing heart failure (HF). However, incident HF is not part of the risk assessment of current guideline-recommended models. The aim of this study was to develop and externally validate the SMART2-HF model for prediction of incident HF in patients with ASCVD. METHODS: SMART2-HF was developed in 7698 individuals with established ASCVD (coronary, cerebrovascular, or peripheral artery disease, or abdominal aortic aneurysm) but without prior HF from the UCC-SMART cohort. Cox proportional hazards models including sex-predictor interactions and with age as the time scale were derived to estimate the 10-year and lifetime risk of incident HF (hospitalization for HF or HF-related death), accounting for competing non-HF mortality. Predictors, limited to routinely available clinical characteristics, were aligned with the SMART2 risk model for recurrent cardiovascular (CV) risk in the same population. External validation was performed in 240 741 patients with ASCVD from six data sources: the Clinical Practice Research Datalink, the HUNT3 study, the SWEDEHEART Registry, the ASCVD-Particles cohort, the Estonian Biobank and the international REACH Registry. RESULTS: During a median follow-up of 11.2 years (interquartile range 6.1-16.4 years), 1031 incident HF events (13%) occurred in the UCC-SMART cohort. In the external validation data sources, a total of 24 885 incident HF events (10%) occurred. The pooled C-statistic was .696 (95% confidence interval .674-.717), with consistent performance in subgroups by sex and type of ASCVD. Predicted risks matched observed incidence in external validation. CONCLUSIONS: The SMART2-HF model enables the prediction of incident HF in patients with ASCVD. Aligned with the guideline-recommended SMART2 model for recurrent CV risk, SMART2-HF can be used as a complementary tool in this population.

Humans↗

Is it clinically possible to distinguish nonhemorrhagic infarct from hemorrhagic stroke?

BACKGROUND AND PURPOSE: Diagnosis of the nonhemorrhagic ischemic type of stroke by analysis of patients' clinical features is considered unreliable because no clinical feature is specific. The diagnosis is so difficult to establish that we cannot hope to use the same method to make a reliable diagnosis in all stroke cases. In this study, we propose a simple scoring system with a positive predictive value of close to 100% to distinguish nonhemorrhagic infarct from hemorrhagic stroke. This scoring is available for all physicians in bedside diagnosis even if this score can be applied to a subgroup of patients. METHODS: Twenty-six clinical variables that might potentially distinguish cerebral hemorrhage from infarction were recorded in patients consecutively admitted to our stroke unit for stroke lasting more than 24 hours with at least unilateral motor weakness affecting face and/or arm and/or leg (internal validity study). Patients previously receiving anticoagulant therapy were excluded. We used CT scan as the gold standard. We used multivariate logistic regression to establish a clinical score from which we derived the classification rule. This rule was validated with data from the next 200 consecutive patients hospitalized in the stroke unit (external validity study). RESULTS: Three hundred sixty-eight patients were enrolled in the internal study. The obtained score was (2 x alcohol consumption) + (1.5 x plantar response) + (3 x headache) + (3 x history of hypertension)--(5 x history of transient neurological deficit)--(2 x peripheral arterial disease)--(1.5 x history of hyperlipidemia)--(2.5 x atrial fibrillation on admission). All patients with a score less than 1 (n = 123) had a nonhemorrhagic infarct (ie, 40% of the 305 patients with a nonhemorrhagic infarct). No threshold was found to diagnose cerebral hemorrhage with a sufficiently high positive predictive value. Among the 200 patients enrolled in the external validity study, 72 patients with a score below 1 had a nonhemorrhagic infarct (ie, 43% of patients with a nonhemorrhagic infarct). CONCLUSIONS: Diagnosis of nonhemorrhagic infarct can be made in 36% (95% confidence interval [CI], 29 to 43) of patients with a high level of accuracy (100% in the external validity study, which gives a 95% CI of 93 to 100). Thus, 43% (95% CI, 36 to 50) of patients with a nonhemorrhagic infarct could receive a bedside diagnosis. The score is simple and can be calculated from information available to all physicians.

Adult↗

Validation of a melanoma prognostic model.

BACKGROUND: A "clinically accessible," 4-variable (patient age, patient sex, tumor location, and tumour thickness) prognostic model has been published previously. This model evaluated variables that were commonly available to the clinician. Because models are heuristic, validity of a prognostic model should be evaluated in a population different from the original population. OBJECTIVE: To evaluate the external validity of this 4-variable melanoma prognostic model. DESIGN: To estimate the external validity of this model, we used a population-based cohort of individuals with melanoma. We also evaluated a 1-variable model (tumor thickness). Estimates of the external validity of these logistic regression models were made using the c statistic and the Brier score. SETTINGS AND PATIENTS: A total of 1261 patients with melanoma evaluated in a multispecialty, university-based practice and 650 patients with melanoma from throughout Connecticut. MAIN OUTCOME MEASURE: Death from melanoma within 5 years of diagnosis. RESULTS: The c statistics for the 4-variable model were 0.86 (95% confidence interval [CI], 0.83-0.89) for the university-based practice data set and 0.81 (95% CI, 0.75-0.86) for the Connecticut data set. For thickness alone, the c statistics were 0.83 (95% CI, 0.80-0.86) and 0.79 (95% CI, 0.74-0.85), respectively. Brier scores for the 4-variable model were 0.09 (95% CI, 0.08-0.10) and 0.08 (95% CI, 0.06-0.09) and for the 1-variable model were 0.09 (95% CI, 0.08-0.10) and 0.08 (95% CI, 0.07-0.10), respectively. No significant differences exist between the data sets for the 4- and 1-variable models. CONCLUSIONS: The 4- and 1-variable models are generalizable. The simpler 1-variable model--tumor thickness--can be used with a relatively small loss in accuracy.

Female↗

Admission of patients with severe and moderate traumatic brain injury to specialized ICU facilities: a search for triage criteria.

OBJECTIVE: To investigate whether triage for direct admission of patients with traumatic brain injury to a trauma center is facilitated by predicting the risk of potentially removable lesions or raised intracranial pressure (ICP). DESIGN AND SETTING: Cohort study in a level I university trauma center. PATIENTS AND PARTICIPANTS: A prospective cohort of primarily (n=200) and secondarily (n=75) referred patients with moderate or severe traumatic brain injury. MEASUREMENTS AND RESULTS: Predictive characteristics for the risk of surgically removable lesions and the risk of raised ICP (repeatedly > or = 20 mmHg) were identified and included in prognostic models. These models were validated internally with bootstrapping techniques and externally on a historic sample (n=205) regarding discriminative ability (AUC). Among the cohort patients, 67% had raised ICP and 54% had surgically removable lesions. Both outcomes occurred more frequently in patients secondarily referred, but the incidence in patients primarily referred was also high (62% and 33% respectively). No strong predictors of raised ICP were identified. Age and pupillary reactivity were significant predictors of surgically removable lesions. The models discriminated reasonably for surgically removable lesions (AUC=0.78 at development and AUC=0.67 at external validation) but not for raised ICP (AUC=0.59 at development and AUC=0.50 at external validation). CONCLUSIONS: It is difficult accurately to identify patients in need of specialized intensive care using baseline characteristics. The high incidence of both outcomes in patients primarily referred support direct admission of more and particularly older patients with severe or moderate brain trauma to level I trauma centers.

Adult↗

Assessment of the validity of a population pharmacokinetic model for epirubicin.

AIMS: The aim of this study was to evaluate a population model for epirubicin clearance using internal and external validation techniques. METHODS: Jackknife samples were used to identify outliers in the population dataset and individuals influencing covariate selection. Sensitivity analyses were performed in which serum aspartate transaminase (AST) values (a covariate in the population model) or epirubicin concentrations were randomly changed by +/-10%. Cross-validation was performed five times, on each occasion using 80% of the data for model development and 20% to assess the performance of the model. External validation was conducted by assessing the ability of the population model to predict concentrations and clearances in a separate group of 79 patients. RESULTS: Structural parameter estimates from all jackknife samples were within 7.5% of the final population estimates and examination of log likelihood values indicated that the selection of AST in the final model was not due to the presence of outliers. Alteration of AST or epirubicin concentrations by +/-10% had a negligible effect on population parameter estimates and their precision. In the cross-validation analysis, the precision of clearance estimates was better in patients with AST concentrations>150 U l-1. In the external validation, epirubicin concentrations were over-predicted by 81.4% using the population model and clearance values were also poorly predicted (imprecision 43%). CONCLUSIONS: The results of internal validation of population pharmacokinetic models should be interpreted with caution, especially when the dataset is relatively small.

Adult↗

Validation of the Rockall risk scoring system in upper gastrointestinal bleeding.

BACKGROUND: Several scoring systems have been developed to predict the risk of rebleeding or death in patients with upper gastrointestinal bleeding (UGIB). These risk scoring systems have not been validated in a new patient population outside the clinical context of the original study. AIMS: To assess internal and external validity of a simple risk scoring system recently developed by Rockall and coworkers. METHODS: Calibration and discrimination were assessed as measures of validity of the scoring system. Internal validity was assessed using an independent, but similar patient sample studied by Rockall and coworkers, after developing the scoring system (Rockall's validation sample). External validity was assessed using patients admitted to several hospitals in Amsterdam (Vreeburg's validation sample). Calibration was evaluated by a chi2 goodness of fit test, and discrimination was evaluated by calculating the area under the receiver operating characteristic (ROC) curve. RESULTS: Calibration indicated a poor fit in both validation samples for the prediction of rebleeding (p<0.0001, Vreeburg; p=0.007, Rockall), but a better fit for the prediction of mortality in both validation samples (p=0.2, Vreeburg; p=0.3, Rockall). The areas under the ROC curves were rather low in both validation samples for the prediction of rebleeding (0.61, Vreeburg; 0.70, Rockall), but higher for the prediction of mortality (0.73, Vreeburg; 0.81, Rockall). CONCLUSIONS: The risk scoring system developed by Rockall and coworkers is a clinically useful scoring system for stratifying patients with acute UGIB into high and low risk categories for mortality. For the prediction of rebleeding, however, the performance of this scoring system was unsatisfactory.

Adolescent↗