PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Internal validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

An observational examination of the literature in diagnostic anatomic pathology.

Original research published in the medical literature confronts the reader with three very basic and closely linked questions--are the authors' conclusions true in the contextual setting in which the work was performed (internally valid); if so, are the conclusions also applicable in other practice settings (externally valid); and, if the conclusions of the study are bona fide, do they represent an important contribution to medical practice or are they true-but-insignificant? Most publications attempt to convince readers that the researchers' conclusions are both internally valid and important, and occasionally papers also directly address external validity. Developing standardized methods to facilitate the prospective determination of research importance would be useful to both journals and their readers, but has proven difficult. In contrast, the evidence-based medicine (EBM) movement has had more success with understanding and codifying factors thought to promote research validity. Of the many variables that can influence research validity, research design is the one that has received the most attention. The present paper reviews the contributions of EBM to understanding research validity, looking for areas where EBM's body of knowledge is applicable to the anatomic pathology (AP) literature. As part of this project, the authors performed a pilot observational analysis of a representative sample of the current pertinent literature on diagnostic tissue pathology. The results of that review showed that most of the latter publications employ one of the four categories of "observational" research design that have been delineated by the EBM movement, and that the most common of these observational designs is a "cross-sectional" comparison. Pathologists do not presently use the "experimental" research designs so admired by advocates of EBM. Slightly > 50% of AP observational studies employed statistical evaluations to support their final conclusions. Comparison of the current AP literature with a selected group of papers published in 1977 shows a discernible change over that period that has affected not just technological procedures, but also research design and use of statistics. Although we feel that advocates of EBM deserve credit for bringing attention to the close link between research design and research validity, much of the EBM effort has centered on refining "experimental" methodology, and the complexities of observational research have often been treated in an inappropriately dismissive manner. For advocates of EBM, an observational study is what you are relegated to as a second choice when you are unable to do an experimental study. The latter viewpoint may be true for evaluating new chemotherapeutic agents, but is unacceptable to pathologists, whose research advances are currently completely dependent on well-conducted observational research. Rather than succumb to randomization envy and accept EBM's assertion that observational research is second best, the challenge to AP is to develop and adhere to standards for observational research that will allow our patients to benefit from the full potential of this time tested approach to developing valid insights into disease.

Anatomy↗

[Ultrasound-controlled anthropometry--on the development of a new method in asymmetry diagnosis].

UNLABELLED: Recently computer-based methods allow the registration of asymmetries in skeleton-axis. In comparison to classical methods the Ultrasound-Guided-Anthropometry (UGA) has been tested in regard to its quality criteria for representing body specific marks in three dimensions. In a sample of healthy subjects objectivity (n = 13), retest-reliability and internal validity (n = 26) have been determined. For analysis of system-quality a standardized cuboid phantom (side length 400 mm) has been measured. RESULTS: objectivity r = 0.93-0.98 (p < .05), retest-reliability r = 0.97-0.99 (p < .05) and internal validity between both methods r = 0.92-0.97 (p < .05). The analysis of system-quality produced an error in measurement of 0.65% (0.58 +/- 1.29 mm). UGA as a screening method is not a substitute for case-history or classical anthropometry, but it offers useful parameters which facilitate decision making for further diagnostic procedures.

Adult↗

Risk factors for injury in child and adolescent sport: a systematic review of the literature.

OBJECTIVE: The objective of this systematic review of the literature is to identify risk factors and potential prevention strategies that may modify risk factors for injury in child and adolescent sport. DATA SOURCES: Seven electronic databases were searched to identify potentially relevant articles. A combination of Medical Subject Headings and text words were used (athletic injuries, sports injury, risk factors, adolescent, and child). STUDY SELECTION: This review is based on epidemiological evidence in which the data are original, an exposure and outcome are objectively measured, and an attempt is made to create a comparison group. Forty-five studies were selected for this review. DATA EXTRACTION: The data summarized include study design, study population, exposures, outcomes, and results. Estimates of odds ratios or relative risks were calculated where study data were adequate to do so. The quality of evidence is based on internal validity, external validity, and causal association. DATA SYNTHESIS: There is some evidence that potentially modifiable risk factors including poor endurance, lack of preseason training, and some psychosocial factors are important risk factors for injury in child and adolescent sport. Concerns with study design, internal validity, and generalizability persist. The evidence is consistent, however, with more convincing evidence from adult population studies. The evidence for nonmodifiable risk factors for injury in adolescent sport (ie, age, sex, previous injury) is consistent among studies. CONCLUSIONS: Sport participation and injury rates in child and adolescent sport are high. This review will assist in targeting the relevant groups and designing future research examining risk factors and prevention strategies in child and adolescent sport. Future clinical trials addressing modifiable risk factors to reduce the incidence of sports injury in this population are necessary.

Accident Prevention↗

Validity of the Stanford-Binet Intelligence Scale-IV: its use in young adults with mental retardation.

The validity of the Stanford Binet-IV (SB-IV) was assessed. This test and the WAIS-R and WRAT-R were administered to 42 adults previously classified with mild to moderate mental retardation. Validity coefficients between scores on the SB-IV and the other two measures were significant. The mean IQ on the SB-IV (mean Test Composite = 43.26) was significantly lower than that on the Wechsler Adult Intelligence Scale-Revised--WAIS-R (mean Full-Scale IQ = 57.91). With regard to the internal validity of the SB-IV, the intersubtest relationships of each of the four Area scores correlated significantly with the Test Composite (range = .66 to .91). Verbal Reasoning earned the highest correlation (.91). Results support the SB-IV's concurrent, criterion-related, and internal validity for use with young adults who have mental retardation.

Adolescent↗

[Analysis of the scientific evidence of the combination therapy in benign prostatic hyperplasia].

OBJECTIVE: The analysis of the scientific evidence of the combination therapy in Benign Prostatic Hyperplasia (BPH). METHODS: 5 published studies about combination therapy in BPH were analysed following the criteria of the Evidence Based Medicine (EBM). Hypothesis, variables, internal validity, results relevance, and the external validity of every study were analysed. RESULTS: Symptoms changes and maximal flow rate (Qmax) improvement were evaluated in four studies and only one analysed the BPH progression. Inclusion and exclusion criteria were similar in the 5 studies. The main reasons of a poor internal validity in 3 studies were the short follow-up, the missing percentage, the absence of a placebo group and a dosage bias. External validity were decreased in the 5 studies by the exclusion criteria and in 3 of them because high doses of alpha-blocker were given to achieve a therapeutic effect. The study with the highest scientific evidence (MTOPS) is the only that offers confidence intervals and number needed to treat. CONCLUSIONS: The Qmax and the symptoms improvement found in all the studies has a moderate clinical relevance. MTOPS study shows that BPH progression has a low incidence that can be highly reduced by means of combination therapy.

Adrenergic alpha-Antagonists↗

Circular instead of hierarchical: methodological principles for the evaluation of complex interventions.

BACKGROUND: The reasoning behind evaluating medical interventions is that a hierarchy of methods exists which successively produce improved and therefore more rigorous evidence based medicine upon which to make clinical decisions. At the foundation of this hierarchy are case studies, retrospective and prospective case series, followed by cohort studies with historical and concomitant non-randomized controls. Open-label randomized controlled studies (RCTs), and finally blinded, placebo-controlled RCTs, which offer most internal validity are considered the most reliable evidence. Rigorous RCTs remove bias. Evidence from RCTs forms the basis of meta-analyses and systematic reviews. This hierarchy, founded on a pharmacological model of therapy, is generalized to other interventions which may be complex and non-pharmacological (healing, acupuncture and surgery). DISCUSSION: The hierarchical model is valid for limited questions of efficacy, for instance for regulatory purposes and newly devised products and pharmacological preparations. It is inadequate for the evaluation of complex interventions such as physiotherapy, surgery and complementary and alternative medicine (CAM). This has to do with the essential tension between internal validity (rigor and the removal of bias) and external validity (generalizability). SUMMARY: Instead of an Evidence Hierarchy, we propose a Circular Model. This would imply a multiplicity of methods, using different designs, counterbalancing their individual strengths and weaknesses to arrive at pragmatic but equally rigorous evidence which would provide significant assistance in clinical and health systems innovation. Such evidence would better inform national health care technology assessment agencies and promote evidence based health reform.

Evidence-Based Medicine↗

Artificial intelligence-assisted histopathological diagnosis of endocervical gastric-type adenocarcinoma: a multicenter model development and validation study.

Endocervical gastric-type adenocarcinoma (GAS) is one of the most aggressive subtypes of cervical cancer and is frequently underdiagnosed due to morphological ambiguity, leading to delayed diagnosis. Despite the availability of molecular and genomic assays, their high cost, complexity, and limited reproducibility restrict clinical use. This study therefore proposes a highly sensitive artificial intelligence (AI)-assisted diagnostic system for GAS based exclusively on H&E-stained histopathological images. We included 309 slides from 96 GAS cases collected at Peking University Third Hospital from January 2018 to January 2025, representing the largest GAS cohort reported to date for AI research. In addition, we incorporated other morphologically analogous diseases, encompassing a total of 1,320 slides sourced from four categories: normal cervical mucosa (NORM), benign endocervical lesion entities (BELE), HPV-associated adenocarcinoma (HPVA), and endometrioid carcinoma with mucinous differentiation (ECMD). We developed GASPath, based on a novel multiple instance learning framework that efficiently captures fine-grained morphological variations from H&E-stained images. Beyond internal validation, GASPath was evaluated across 12 independent retrospective cohorts and further subjected to large-scale real-world validation on more than 7,000 samples from March 2024 to April 2025. Across three stages, GASPath demonstrated high performance. In internal validation (Stage I), it achieved an accuracy of 0.980 (95% CI 0.977-0.983) and an ROC-AUC of 0.995 (95% CI 0.994-0.997). In external validation (Stage II), the sensitivity reached 0.902 and improved to 0.968 with proposed strategies. For biopsy samples, GASPath achieved an ROC-AUC of 0.990 (95% CI 0.984-0.997). In large-scale real-world deployment (Stage III, n&#x2009;=&#x2009;7,056), GASPath achieved a balanced accuracy of 0.953, with 100% sensitivity for GAS (45/45 cases correctly identified). The heatmaps highlight morphological features of GAS that are easily underestimated, such as irregular, angulated glands, subtle loss of nuclear polarity, and mild cytologic atypia, which show substantial morphological overlap with other diagnostic categories. GASPath enables high-sensitivity detection of GAS in routine H&E-stained slides, obviating the need for extensive auxiliary testing while preventing underdiagnosis and misdiagnosis. This advancement addresses a critical gap by streamlining diagnostic workflows without compromising accuracy. Its implementation could enable cost-effective, scalable AI-assisted diagnostics, potentially transforming the early detection and management of this aggressive cancer subtype.

Female↗

Validity and internal consistency of a whiplash-specific disability measure.

STUDY DESIGN: Cross-sectional study of patients with whiplash-associated disorders investigating the internal consistency, factor structure, response rates, and presence of floor and ceiling effects of the Whiplash Disability Questionnaire (WDQ). OBJECTIVES: The aim of this study was to confirm the appropriateness of the proposed WDQ items. SUMMARY OF BACKGROUND DATA: Whiplash injuries are a common cause of pain and disability after motor vehicle accidents. Neck disability questionnaires are often used in whiplash studies to assess neck pain but lack content validity for patients with whiplash-associated disorders. The newly developed WDQ measures functional limitations associated with whiplash injury and was designed after interviews with 83 patients with whiplash in a previous study. METHODS: Researchers sought expert opinion on items of the WDQ, and items were then tested on a clinical whiplash population. Data were inspected to determine floor and ceiling effects, response rates, factor structure, and internal consistency. Packages of questionnaires were distributed to 55 clinicians, whose patients with whiplash completed and returned 101 questionnaires to researchers. RESULTS: No substantial floor or ceiling effects were identified on inspection of data. The overall floor effect was 12%, and the overall ceiling effect was 4%. Principal component analysis identified one broad factor that accounted for 65% of the variance in responses. Internal consistency was high; Cronbach's alpha = 0.96. CONCLUSIONS: Results of the study supported the retention of the 13 proposed items in a whiplash-specific disability questionnaire. Dependent on the results of further psychometric testing, the WDQ is likely to be an appropriate outcome measure for patients with whiplash.

Adult↗

Validation of the SADL questionnaire.

OBJECTIVE: To cross-validate the psychometric characteristics of the Satisfaction with Amplification in Daily Life (SADL) questionnaire (Cox & Alexander, 1999), and to explore the SADL's construct validity. DESIGN: Thirteen private practice Audiology clinics each distributed SADL questionnaires, by mail, to 20 adults who had recently obtained hearing aids. The completed questionnaires were returned to a central site and subject anonymity was assured. There were 196 usable responses. RESULTS: Psychometric characteristics of the items were found to be very similar to those reported previously. Thus, the internal validity of the instrument was strongly supported. The assumption that the SADL quantifies satisfaction by assessing its components was evaluated by examining the relationship between SADL scores and scores on a traditional single-item satisfaction measure. A logical and statistically significant relationship was seen between the two measures, thereby supporting the construct validity of both types of data. For private-pay clients, satisfaction scores were very similar to the interim norms published by Cox and Alexander (1999). However, clients whose hearing aids were partly or fully purchased by insurance or benefits programs tended to be more satisfied than interim norms for third-party pay clients derived 5 yr ago. For most types of clients, there was a tendency toward more satisfaction in the Negative Features subscale than observed in our previous research. CONCLUSIONS: Both construct and internal validity of the SADL questionnaire were supported by this research. The previously published interim norms appear to be mostly appropriate for private-pay clients, but might require adjustment in the Negative Features subscale. Further research is needed to explore the relationship between satisfaction and device purchase issues (third-party versus private pay).

Adult↗

Identifying subgroups among poor prognosis patients with nonseminomatous germ cell cancer by tree modelling: a validation study.

BACKGROUND: In order to target intensive treatment strategies for poor prognosis patients with non-seminomatous germ cell cancer, those with the poorest prognosis should be identified. These patients might profit most from more intensive treatment strategies. For this purpose, a regression tree was previously developed on 332 patients. We aimed to evaluate the performance and structure of this tree. PATIENTS AND METHODS: The previously developed tree was applied to 456 patients with a poor prognosis as defined by the International Germ Cell Cancer Collaborative Group (IGCCCG). Next, we developed a new tree to evaluate whether a similar structure to the previous tree was found. We assessed the internal validity of the new tree, and compared the 2-year survival estimates of each subgroup together with the discriminative ability for both the previously developed and the new tree. Discriminative ability was measured by a concordance (c) statistic, which varies between 0.5 (no discrimination) and 1.0 (perfect discrimination). RESULTS: The 2-year survival estimates in the IGCCCG data ranged from 33% to 63%. The ordering of the subgroups was different and discriminative ability was lower than originally found (c = 0.56 in the IGCCCG data versus 0.63 originally). The new tree differed considerably from the original tree, and identified poor prognosis subgroups with 2-year survival estimates from 38% to 73%. Internal validation showed similar discriminative ability for the new tree and the original tree (c = 0.59 versus 0.56). CONCLUSIONS: The previously developed tree showed poor validity with respect to discriminative ability and the stability of its structure. The performance of the new tree was also unsatisfactory. Given the low proportion of patients categorised as poor prognosis, it seems that the potential to identify further subgroups with the currently available patient characteristics is limited.

Adolescent↗

Sentence completion test for depression (SCD): an idiographic measure of depressive thinking.

OBJECTIVES: This study set out to investigate the reliability and validity of the Sentence Completion Test for Depression (SCD) as a clinical measure. In contrast to questionnaire measures of depressive thinking, respondents finish incomplete sentences using their own words. This elicits idiographic information concurrent with measuring depressive thinking. METHOD: In Study 1, measures of negative thinking were tested between a depressed group and a non-depressed control group. A preliminary item analysis was conducted and replicated on separate samples in Study 2. Psychometric properties of the test were investigated. In Study 3, idiographic validity and sensitivity to change were explored in a sample of clinical cases with reference to cognitive-behavioural case-formulation. RESULTS: In Study 1, the depressed group produced more negatives and fewer positives, and the SCD demonstrated good content validity, internal consistency and inter-rater reliability. The preliminary short-form had comparable psychometric properties, and these were replicated on new samples in Study 2. Sensitivity and specificity values were above 90% in both studies. In Study 3, idiographic content generated hypotheses about target problems and dysfunctional beliefs within cognitive-behavioural case-formulation, and SCD scores were sensitive to clinical change. CONCLUSIONS: The SCD demonstrates good construct validity, internal consistency, inter-rater reliability, sensitivity, and specificity. It offers an idiographic assessment of depression that is complementary to questionnaire measures, particularly by generating hypotheses about target problems and dysfunctional beliefs within a cognitive-behavioural case-formulation. This is achieved without loss to reliability and validity at the nomothetic level.

Adult↗

Corticosteroid injections for lateral epicondylitis: a systematic review.

Patients with lateral epicondylitis (tennis elbow) are frequently treated with corticosteroid injections, in order to relieve pain and diminish disability. The objective of this review was to evaluate the effectiveness of corticosteroid injections for lateral epicondylitis. Randomised controlled trials (RCTs) were identified by a highly sensitive search strategy in six databases in combination with reference tracking. Two independent reviewers selected and assessed the methodological quality of RCTs that included patients with lateral epicondylitis treated with corticosteroid injection(s), and reported at least one clinically relevant outcome measure. Standardised mean differences were computed for continuous data and relative risks (RR) for dichotomous data. A best-evidence synthesis was conducted, weighting the studies with respect to their internal validity, statistical significance, clinical relevance, and statistical power. Thirteen studies consisting of 15 comparisons were included in the review, evaluating the effects of corticosteroid injections compared to placebo injection (n=2), injection with local anaesthetic (n=5), another conservative treatment (n=5), or another corticosteroid injection (n=3). Almost all studies had poor internal validity scores. For short-term outcomes ( or=6 months), no statistically significant or clinically relevant results in favour of corticosteroid injections were found. Although the available evidence shows superior short-term effects of corticosteroid injections for lateral epicondylitis, it is not possible to draw firm conclusions on the effectiveness of injections, due to the lack of high quality studies. No beneficial effects were found for intermediate or long-term follow-up. More, better designed, conducted and reported RCTs with intermediate and long-term follow-up are needed.

Adrenal Cortex Hormones↗

Towards measurement of outcome for patients with varicose veins.

OBJECTIVE: To develop a valid and reliable outcome measure for patients with varicose veins. DESIGN: Postal questionnaire survey of patients with varicose veins. SETTING: Surgical outpatient departments and training general practices in Grampian region. SUBJECTS: 373 patients, 287 of whom had just been referred to hospital for their varicose veins and 86 who had just consulted a general practitioner for this condition and, for comparison, a random sample of 900 members of the general population. MAIN MEASURES: Content validity, internal consistency, and criterion validity. RESULTS: 281(76%) patients (mean age 45.8; 76% female) and 542(60%) of the general population (mean age 47.9; 54% female) responded. The questionnaire had good internal consistency as measured by item-total correlations. Factor analysis identified four important health factors: pain and dysfunction, cosmetic appearance, extent of varicosity and complications. The validity of the questionnaire was demonstrated by a high correlation with the SF-36 health profile, which is a general measure of patients' health. The perceived health of patients with varicose veins, as measured by the SF-36, was significantly lower than that of the sample of the general population adjusted for age and a lower proportion of women. CONCLUSION: A clinically derived questionnaire can provide a valid and reliable tool to assess the perceived health of patients with varicose veins. IMPLICATIONS: The questionnaire may be used to justify surgical treatment of varicose veins.

Adult↗

Botulinum toxin type A therapy for blepharospasm.

BACKGROUND: Blepharospasm is a focal dystonia characterized by chronic intermittent or persistent involuntary eyelid closure due to spasmodic contractions of the orbicularis oculi muscles. Other facial and neck muscles are also frequently involved. Most cases are idiopathic and blepharospasm is generally a life-long disorder. Its severity can range from repeated frequent blinking to persistent forceful closure of the eyelids with functional blindness. Botulinum toxin type A (BtA) is the current first line therapy. OBJECTIVES: To determine whether botulinum toxin (BtA) is an effective and safe treatment for blepharospasm. SEARCH STRATEGY: We identified studies for inclusion in the review using the Cochrane Movement Disorders Group trials register, the Cochrane Central Register of Controlled Trials (CENTRAL), MEDLINE, EMBASE, handsearches of the Movement Disorders Journal and abstracts of international congresses on movement disorders and botulinum toxin, communication with other researchers in the field, reference lists of papers found using above search strategies, and contact with authors and drug manufacturers. SELECTION CRITERIA: Studies were eligible for inclusion in the review if they evaluated the efficacy of BtA for the treatment of blepharospasm. They must have been randomised and placebo-controlled. DATA COLLECTION AND ANALYSIS: We used a paper pro-forma to collect data from the included studies using double extraction by two independent reviewers. The two reviewers separately assessed each trial for internal validity and they settled differences between them by discussion. The outcome measures used included adverse events, improvement in symptomatic rating scales, subjective evaluation by patients and clinicians, and changes in quality of life assessments. MAIN RESULTS: We found few controlled trials. They were of short duration and enrolled small numbers of patients. Because of their poor internal validity, the characteristics of the populations studied, and the types of interventions and outcomes, none of the trials fitted our criteria for inclusion. However, all these trials found BtA to be superior to placebo as did large case-control and cohort studies, which reported that around 90% of patients benefited. AUTHORS' CONCLUSIONS: There are no high quality, randomised, controlled efficacy data to support the use of Bt for blepharospasm. Despite this, other studies suggest that BtA is highly effective and safe for treating blepharospasm and support its use. The effect size (90% of patients benefit) seen in open studies makes it very difficult and probably unethical to perform new placebo-controlled trials of efficacy of BtA for blepharospasm. Future trials should explore technical factors such as the optimum treatment intervals, different injection techniques, doses, Bt types and formulations. Other issues include service delivery, quality of life, long-term efficacy, safety, and immunogenicity.

Blepharospasm↗

Effectiveness of exercise therapy and manual mobilisation in ankle sprain and functional instability: a systematic review.

This study critically reviews the effectiveness of exercise therapy and manual mobilisation in acute ankle sprains and functional instability by conducting a systematic review of randomised controlled trials. Trials were searched electronically and manually from 1966 to March 2005. Randomised controlled trials that evaluated exercise therapy or manual mobilisation of the ankle joint with at least one clinically relevant outcome measure were included. Internal validity of the studies was independently assessed by two reviewers. When applicable, relative risk (RR) or standardised mean differences (SMD) were calculated for individual and pooled data. In total 17 studies were included. In thirteen studies the intervention included exercise therapy and in four studies the effects of manual mobilisation of the ankle joint was evaluated. Average internal validity score of the studies was 3.1 (range 1 to 7) on a 10-point scale. Exercise therapy was effective in reducing the risk of recurrent sprains after acute ankle sprain: RR 0.37 (95% CI 0.18 to 0.74), and with functional instability: RR 0.38 (95% CI 0.23 to 0.62). No effects of exercise therapy were found on postural sway in patients with functional instability: SMD: 0.38 (95% CI -0.15 to 0.91). Four studies demonstrated an initial positive effect of different modes of manual mobilisation on dorsiflexion range of motion. It is likely that exercise therapy, including the use of a wobble board, is effective in the prevention of recurrent ankle sprains. Manual mobilisation has an (initial) effect on dorsiflexion range of motion, but the clinical relevance of these findings for physiotherapy practice may be limited.

Acute Disease↗

A meta-evaluation of smoking cessation intervention research among pregnant women: improving the science and art.

In 1986 Windsor and Orleans described guidelines and standards to evaluate the quality of smoking cessation intervention research among pregnant women. This paper presents a meta-evaluation (ME) of the evaluation research in this area from 1986 to 1998. ME is defined as a systematic review of experimental and quasi-experimental evaluation research using a standardized set of methodological criteria to rate the internal validity--efficacy or effectiveness--of intervention results. Five criteria were used to rate 23 smoking cessation intervention studies among pregnant smokers in prenatal care: (1) evaluation research design, (2) sample representativeness, sample size and power estimation, (3) population characteristics, (4) measurement quality, and (5) replicability of interventions. Eleven studies had sufficient methodological quality to produce results of high internal validity. Poor measurement of smoking status, patient selection biases and incorrect calculation of quit rates were the major methodological weakness. Recommendations for future evaluation research are made.

Female↗

Major adverse outcomes after percutaneous transluminal coronary angioplasty: a clinical prediction rule.

In this study, we developed and internally validated a clinical model for predicting major adverse outcomes in patients undergoing percutaneous transluminal coronary angioplasty (PTCA) using a multi-institutional prospective cohort study involving all adult patients who underwent PTCA at 12 participating institutions from August 1993 to October 1995. A major adverse outcome, defined as death, renal failure, myocardial infarction, cardiac arrest, stroke, or coma, occurred in 3.3 and 3.2% of patients in the derivation and validation sets, respectively. Death occurred in 1.5% in both sets. Fourteen variables were independently correlated with major adverse outcomes. The rule, which stratifies PTCA patients into six levels of risk based on the severity score, showed excellent discrimination (receiver-operating characteristic curve area 0.82) and calibration (Hosmer-Lemeshow chi-square statistic P =.90) and performed well on internal validation. This rule allows accurate preprocedure stratification of PTCA candidates according to their risk of suffering a major adverse outcome.

Angioplasty, Balloon, Coronary↗

Artificial neural network is superior to MELD in predicting mortality of patients with end-stage liver disease.

BACKGROUND: Despite its accuracy, the model for end-stage liver disease (MELD), currently adopted to determine the prognosis of patients with liver cirrhosis, guide referral to transplant programmes and prioritise the allocation of donor organs, fails to predict mortality in a considerable proportion of patients. AIMS: To evaluate the possibility to better predict 3-month liver disease-related mortality of patients awaiting liver transplantation using an artificial neural network (ANN). PATIENTS AND METHODS: The ANN was constructed using data from 251 consecutive people with cirrhosis listed for liver transplantation at the Liver Transplant Unit, Bologna, Italy. The ANN was trained to predict 3-month survival on 188 patients, tested on the remaining 63 (internal validation group) unknown by the system and finally on 137 patients listed for liver transplantation at the King's College Hospital, London, UK (external cohort). Predictions of survival obtained with ANN and MELD on the same datasets were compared using areas under receiver-operating characteristic (ROC) curves (AUC). RESULTS: The ANN performed significantly better than MELD both in the internal validation group (AUC = 0.95 v 0.85; p = 0.032) and in the external cohort (AUC = 0.96 v 0.86; p = 0.044). CONCLUSIONS: The ANN measured the mortality risk of patients with cirrhosis more accurately than MELD and could better prioritise liver transplant candidates, thus reducing mortality in the waiting list.

Area Under Curve↗