PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “External validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Machine learning-enabled multi-omics discovery of prognostic biomarkers and signaling targets in pancreatic cancer.

Pancreatic ductal adenocarcinoma (PDAC) remains difficult to subtype using single omics layers. We conducted an exploratory investigation integrating reverse-phase protein array (RPPA) and DNA methylation data from the cancer genome atlas (TCGA)- pancreatic adenocarcinoma (PAAD) to assess the feasibility of multi-omics subtyping, alongside a supervised machine learning analysis of a small gene expression omnibus (GEO) transcriptomic cohort (n = 26) to identify candidate diagnostic genes. RPPA-based K-means clustering suggested a weak, possible two-subtype structure (silhouette ≈ 0.16) that remained unassociated with overall survival (log-rank p = 0.113) and lacked independent prognostic value. An independently performed similarity network fusion (SNF) analysis integrating RPPA and methylation data showed low concordance with RPPA-derived subtypes (Adjusted Rand Index (ARI) = 0.014), indicating limited convergence between molecular modalities. Supervised machine learning analysis of the GEO cohort using a fully nested leave-one-out cross-validation pipeline achieved a mean (area under the curve) AUC of 0.896 across four classifiers and identified four-fold-stable candidate genes (ESCO2, COL17A1, BCL2L14, and SOWAHB). However, this gene panel demonstrated limited external validity across two independent PDAC cohorts (log-rank p = 0.438 for both GSE62452 and GSE28735), indicating limited generalizability despite robust internal performance. Collectively, these findings provide limited evidence for a robust, prognostically significant multi-omics subtype or a validated diagnostic gene signature; instead, this study serves as a hypothesis-generating resource and highlights the importance of rigorous cross-validation and independent external validation in small-sample transcriptomic biomarker discovery.

Humans↗

Pragmatic controlled clinical trials in primary care: the struggle between external and internal validity.

BACKGROUND: Controlled clinical trials of health care interventions are either explanatory or pragmatic. Explanatory trials test whether an intervention is efficacious; that is, whether it can have a beneficial effect in an ideal situation. Pragmatic trials measure effectiveness; they measure the degree of beneficial effect in real clinical practice. In pragmatic trials, a balance between external validity (generalizability of the results) and internal validity (reliability or accuracy of the results) needs to be achieved. The explanatory trial seeks to maximize the internal validity by assuring rigorous control of all variables other than the intervention. The pragmatic trial seeks to maximize external validity to ensure that the results can be generalized. However the danger of pragmatic trials is that internal validity may be overly compromised in the effort to ensure generalizability. We are conducting two pragmatic randomized controlled trials on interventions in the management of hypertension in primary care. We describe the design of the trials and the steps taken to deal with the competing demands of external and internal validity. DISCUSSION: External validity is maximized by having few exclusion criteria and by allowing flexibility in the interpretation of the intervention and in management decisions. Internal validity is maximized by decreasing contamination bias through cluster randomization, and decreasing observer and assessment bias, in these non-blinded trials, through baseline data collection prior to randomization, automating the outcomes assessment with 24 hour ambulatory blood pressure monitors, and blinding the data analysis. SUMMARY: Clinical trials conducted in community practices present investigators with difficult methodological choices related to maintaining a balance between internal validity (reliability of the results) and external validity (generalizability). The attempt to achieve methodological purity can result in clinically meaningless results, while attempting to achieve full generalizability can result in invalid and unreliable results. Achieving a creative tension between the two is crucial.

Blood Pressure↗

Quantification of communication processes, is it possible?

The PAS system (Problem-Analysis-Solution-system) is developed to quantify oral communication processes during counselling in pharmacy practice. The pharmacist translates the patient's drug-related questions into a P-code, the analysis of the question into an A-code and finally the given solution upon the question into a S-code. The PAS system has been developed for two goals. First, for the registation of drug-related questions from patients which gives the pharmacist insight in the most common issues addressed by patients. Second, it might help the pharmacist to structure the communication with the patient during the consultation. Forty-one pharmacists participated in the evaluation of the PAS system. The validation of the PAS system consisted of two phases: the external validation and the internal validation. Kappa values were calculated as a measure of agreement in the coding by the pharmacists. The kappa-value of the external validation for the P-, A- and S-codes for the total set of questions indicate a moderate to poor agreement. This means that pharmacists categorize drug-related questions from patients in a different way. Therefore we conclude that the PAS system is less reliable for research purpose. The kappa-value of the internal validation for the P-code varies from 0.42 to 0.91. For the A-code it varies from 0.07 to 0.35 and for the S-code from zero to 0.68. Internal reproducibility is good for P-code but not for the A-code and S-code. This implies that the pharmacist can use the P-codes for registration of patients' questions in his own pharmacy. Moreover, the usage of the PAS system during counselling in pharmacy practice can structure the consultation.

Communication↗

Treatment satisfaction of patients with lower urinary tract symptoms: randomised controlled trials vs. real life practice.

Randomised controlled trials (RCTs) are an important scientific tool to determine the efficacy and tolerability of a given treatment relative to placebo or other treatment forms. However, due to strict inclusion and exclusion criteria the patient populations in RCTs may not be fully representative for those routinely consulting the physician. Moreover, participation in a formal study puts physician and patient in a situation where they may react different than in real life. In contrast real life practice (RLP) studies cannot determine treatment efficacy or tolerability in absolute terms since they typically do not include a control group and are purely observational. On the other hand, they tend to be more representative for real treatment outcomes. Thus, RCTs have high internal but less external validity whereas RLP studies have less internal and greater external validity. Hence, RCTs and RLP studies should not be considered as mutually exclusive but rather as complementing each other. Specific advantages and disadvantages of RCTs and RLP studies will be discussed using published evidence for the treatment of lower urinary tract symptoms suggestive of benign prostatic obstruction with alpha1-adrenoceptor antagonists and other treatments.

Humans↗

A clinical and echocardiographic score for assigning risk of major events after dobutamine echocardiograms.

OBJECTIVES: We sought to develop and validate a risk score combining both clinical and dobutamine echocardiographic (DbE) features in 4890 patients who underwent DbE at three expert laboratories and were followed for death or myocardial infarction for up to five years. BACKGROUND: In contrast to exercise scores, no score exists to combine clinical, stress, and echocardiographic findings with DbE. METHODS: Dobutamine echocardiography was performed for evaluation of known or suspected coronary artery disease in 3156 patients at two sites in the U.S. After exclusion of patients with incomplete follow-up, 1456 DbEs were randomly selected to develop a multivariate model for prediction of events. After simplification of each model for clinical use, the models were internally validated in the remaining DbE patients in the same series and externally validated in 1733 patients in an independent series. RESULTS: The following score was derived from regression models in the modeling group (160 events): DbE risk = (age.0.02) + (heart failure + rate-pressure product <15000).0.4 + (ischemia + scar).0.6. The presence of each variable was scored as 1 and its absence scored as 0, except for age (continuous variable). Using cutoff values of 1.2 and 2.6, patients were classified into groups with five-year event-free survivals >95%, 75% to 95%, and <75%. Application of the score in the internal validation group (265 events) gave equivalent results, as did its application in the external validation group (494 events, C index = 0.72). CONCLUSIONS: A risk score based on clinical and echocardiographic data may be used to quantify the risk of events in patients undergoing DbE.

Cardiotonic Agents↗

Cost estimates for hospital inpatient care in Australia: evaluation of alternative sources.

OBJECTIVE: This paper presents a framework for evaluation of alternative sources of estimates of the costs of hospital inpatient care in Australia. It argues that the choice of costing methods depends on the decision-context and the sensitivity of the decision to estimation errors. METHOD: Five criteria are proposed for evaluation of sources of hospital cost data, with detailed consideration of the way estimates are derived in two computerised approaches which use accounting data. Three broad approaches to cost estimation are evaluated against these criteria. RESULTS: Choosing an estimation method entails an optimisation analysis for each decision context. 'Microcosting' techniques remains the most valid approach to cost estimation, but are costly and this may, in turn, limit the sample of patients or institutions. Protocol-based cost estimates vary widely in their validity, depending on source data, but there is little justification for continued use of crude per diem cost estimates in such protocols. When precision and resolution are important objectives, clinical costing approaches provide the most valid inpatient cost estimates at a reasonable data cost. When external validity is important, or where standardisation of hospital costs is desired, use of published national cost weights may be preferred. CONCLUSION: Both primary and secondary sources of cost data must withstand challenges to internal and external validity. The 'resolution' (or precision) of cost estimates and the relative costs of collection must also be considered. IMPLICATIONS: Studies using estimates of the costs of hospital care should defend the appropriateness of the costing approach and data source for the decision context.

Accounting↗

Treatment research at the crossroads: the scientific interface of clinical trials and effectiveness research.

OBJECTIVE: Policy and clinical management decisions depend on data on the health and cost impacts of psychiatric treatments under usual care, i.e., effectiveness. Clinical trials, however, provide information on treatment efficacy under best-practice conditions. An understanding of the design, analysis, and conventions of both efficacy and effectiveness studies can lead to research that better informs clinical and societal questions. METHOD: This paper contrasts the strengths and limitations of clinical trials and effectiveness studies for addressing policy and clinical decisions. These research approaches are assessed in terms of outcomes, treatments, service delivery context, implementation conventions, and validity. RESULTS: Clinical trials and effectiveness research share problems of internal and external validity despite more attention to internal validity in clinical trials (e.g., randomization, blinding, standardized protocols) and to external validity in effectiveness studies (e.g., community-based treatments, representative samples). CONCLUSIONS: To develop research at the interface of clinical trials and effectiveness studies, research goals must be redefined, and methods, such as cost-utility and econometric analyses, must be shared and developed. Development of hybrid designs that combine features of efficacy and effectiveness research will require separation of conventions such as frequency of follow-up, intensity of measurement, and sample size from the central scientific issues of aims and validity.

Clinical Protocols↗

The Bech-Rafaelsen Melancholia Scale (MES) in clinical trials of therapies in depressive disorders: a 20-year review of its use as outcome measure.

OBJECTIVE: To evaluate the psychometric properties of the Bech-Rafaelsen Melancholia Scale (MES) by reviewing clinical trials in which it has been used as outcome measure. METHOD: The psychometric analysis included internal validity (total scores being a sufficient statistic), interobserver reliability, and external validity (responsiveness in short-term trials and relapse prevention in long-term trials). RESULTS: The results showed that the MES is a unidimensional scale, indicating that the total score is a sufficient statistic. The interobserver reliability of the MES has been found adequate both in unipolar and bipolar depression. External validity including both relapse, response and recurrence indicated that the MES has a high responsiveness and sensitivity. CONCLUSION: The MES has been found a valid and reliable scale for the measurement of changes in depressive states during short-term as well as long-term treatment.

Clinical Trials as Topic↗

Pretreatment nomogram that predicts 5-year probability of metastasis following three-dimensional conformal radiation therapy for localized prostate cancer.

PURPOSE: There are several nomograms for the patient considering radiation therapy for clinically localized prostate cancer. Because of the questionable clinical implications of prostate-specific antigen (PSA) recurrence, its use as an end point has been criticized in several of these nomograms. The goal of this study was to create and to externally validate a nomogram for predicting the probability that a patient will develop metastasis within 5 years after three-dimensional conformal radiation therapy (CRT). PATIENTS AND METHODS: We conducted a retrospective, nonrandomized analysis of 1,677 patients treated with three-dimensional CRT at Memorial Sloan-Kettering Cancer Center (MSKCC) from 1988 to 2000. Clinical parameters examined were pretreatment PSA level, clinical stage, and biopsy Gleason sum. Patients were followed until their deaths, and the time at which they developed metastasis was noted. A nomogram for predicting the 5-year probability of developing metastasis was constructed from the MSKCC cohort and validated using the Cleveland Clinic series of 1,626 patients. RESULTS: After three-dimensional CRT, 159 patients developed metastasis. At 5 years, 11% of patients experienced metastasis by cumulative incidence analysis (95% CI, 9% to 13%). A nomogram constructed from the data gathered from these men showed an excellent ability to discriminate among patients in an external validation data set, as shown by a concordance index of 0.81. CONCLUSION: A nomogram with reasonable accuracy and discrimination has been constructed and validated using an external data set to predict the probability that a patient will experience metastasis within 5 years after three-dimensional CRT.

Adenocarcinoma↗

Measurement of depression in patients with chronic obstructive pulmonary disease (COPD).

OBJECTIVE: To estimate the validity of the Hamilton Depression Scale (HDS) in a population of patients with chronic obstructive pulmonary disease (COPD). METHODS: Forty-nine patients with moderate to severe COPD were examined using the ICD-10 criteria for depression. The mean age of the patients was 71 years and 33 (64%) were women. Forty-six (94%) of the patients were also evaluated using the 17-item HDS including the six-item Hamilton Depression subscale (HDSS). Internal and external validity were measured using factor analysis, Cronbach Coefficient alpha, Loevinger coefficient of homogeneity, correlation analysis and ROC-curves. RESULTS: Twenty-three (47%) of the patients were depressed according to the ICD-10 criteria for depression. The HDSS but not the HDS showed a good internal validity. An acceptable external validity was furthermore shown for the HDSS. CONCLUSION: The HDSS can be recommended as a suitable depression rating scale for COPD patients.

Aged↗

Identifying patients for blood conservation strategies.

BACKGROUND: Generally, only the type of operation is used to estimate the need for perioperative homologous blood transfusion. This study quantified the extent to which the estimation could be improved if, in addition, simple patient characteristics were taken into account. METHODS: Retrospective data on 24 509 consecutive adult surgical patients were used to derive and validate three models to predict perioperative homologous transfusion. The first model was a univariable model with type of operation as the only predictor. The second and third models were a full and a simplified multivariable logistic regression model. The performance of the multivariable models was tested in two validation sets: in similar patients who had operations in the same general hospital (internal validation) and in patients who had operations in a university hospital (external validation). The areas under the receiver-operator characteristic (ROC) curve were compared with that found in the derivation set. RESULTS: There were no important differences in characteristics between the derivation and validation sets. The ROC area of the model including surgery only was 0.92 (99 per cent confidence interval (c.i.) 0.91 to 0.94) and that of the full and simplified multivariable models 0.95 (99 per cent c.i. 0.94 to 0.96) and 0.94 (99 per cent c.i. 0.93 to 0.95) respectively. The latter two were significantly different from the first one. In the external validation set the ROC area of the simplified model was 0.84 (95 per cent c.i. 0.83 to 0.86). Patients who had a preoperative haemoglobin level lower than 13 g/dl and underwent major invasive surgery had the highest risk (43 per cent) of transfusion. CONCLUSION: A simple algorithm using type of operation and haemoglobin concentration was effective in identifying patients likely to need perioperative homologous blood transfusion.

Adult↗

Assessing socioeconomic status in adolescents: the validity of a home affluence scale.

STUDY OBJECTIVE: To examine the completion rate, internal reliability, and external validity of a home affluence scale based on adolescents' reports of material circumstances in the home as a measure of family socioeconomic status. DESIGN: Cross sectional survey. SETTING: Data were collected from a school based study in seven schools in the north of England Cheshire over a five month period from September 1999 to January 2000. PARTICIPANTS: 1824 students (1248 girls, 567 boys) aged 13-15 years who were attending normal classes in Years 9 and 10 in 7 schools on the days of data collection. MAIN RESULTS: Comparatively poor completion rates were found for questions on parental education and occupation while material deprivation items had much higher completion rates. There was evidence that students with poorer material circumstances were less able to report parental education and occupation whereas material based questions showed less bias. A home affluence scale composed of material items was found to have adequate internal reliability and good external validity. CONCLUSIONS: A home affluence scale based on material markers provides a useful alternative in assessing family affluence in adolescents. Additionally, it prevents exclusion of those less materially well off adolescents who fail to complete conventional socioeconomic status items.

Adolescent↗

Evaluation and application of models for the prediction of ready biodegradability in the MITI-I test.

Three existing models and one newly developed model for the prediction of ready biodegradability of organic compounds are evaluated by comparing the descriptors they use, and the consistency of the models when applied to the set of High Production Volume Chemicals (HPVC) in the European Union. Linear regression models developed for the OECD showed the best performance in the external validation (84.7% correct), although comparison with the other three models is flawed because of the class specificity of these models. With these models 567 of the 894 compounds could be predicted in the validation. The multivariate statistical model showed the best performance in the external validation (82.7% correct) combined with the broadest applicability of the model. The evaluation of the predictions of the models for the HPVC shows that all models are highly consistent in their prediction of not-ready biodegradability, but much less consistency is seen in the prediction of ready biodegradability. This complies with the observation that all 4 models show better performance in their predictions of not-ready biodegradability.

Analysis of Variance↗

Validation of a tool to safely triage selected patients with chest pain to unmonitored beds.

OBJECTIVE: To externally validate a chest pain protocol that triages low risk patients with chest pain to an unmonitored bed. METHODS: Retrospective study of all patients admitted from the emergency department of a tertiary referral public teaching hospital with an admission diagnosis of 'unstable angina' or suspected ischemic chest pain. Data was collected on adverse outcomes and analysed on the basis of intention-to-treat according to the chest pain protocol. RESULTS: There were no life-threatening arrhythmias, cardiac arrests or deaths within the first 72 h of admission in the group assigned to an unmonitored bed by the chest pain protocol ([0/244]; 0.0%: 95% confidence interval 0.0-1.5%). Four patients had an uncomplicated myocardial infarction, two patients had recurrent ischemic chest pain and one patient developed acute pulmonary oedema ([7/244]; 2.9%: 95% confidence interval 1.2-5.8%). CONCLUSION: This retrospective study externally validated the chest pain protocol. Care in a monitored bed would not have altered outcomes for patients triaged to an unmonitored bed by the chest pain protocol. Compared to current guidelines, application of the chest pain protocol could increase the availability of monitored beds.

Aged↗

Randomized and non-randomized patients in clinical trials: experiences with comprehensive cohort studies.

In clinical research, randomized trials are widely accepted as the definitive method of evaluating the efficacy of therapies. Random assignment of patients to treatment ensures internal validity of the comparison of new treatments with controls. An assessment of external validity can best be achieved by comparing the randomized study sample to the population of patients who met the eligibility criteria but did not consent to randomization. The Comprehensive Cohort Study (CCS) is designed to recruit all patients fulfilling the clinical eligibility criteria regardless of their consent to randomization. The CCS concept was adopted in the major clinical trials of the German Breast Cancer Study Group (GBSG) conducted between 1983 and 1989. In this period 124 centres recruited 2084 patients in three clinical trials. 734 (35 per cent) of these patients accepted being randomized, while 1350 (65 per cent) chose one of the treatments under study; the randomization rates differed remarkably between trials. In this paper we examine the representativeness of the randomized patients in the three trials. Based on a median follow-up of about 5 years we present results on the external validity of the treatment effects estimated in the randomized patients by means of Cox's proportional hazards model and compare them between trials. We discuss advantages and disadvantages of the CCS design and conclude that its use is only justified under extraordinary circumstances.

Breast Neoplasms↗

Pros and cons of permutation tests in clinical trials.

Hypothesis testing, in which the null hypothesis specifies no difference between treatment groups, is an important tool in the assessment of new medical interventions. For randomized clinical trials, permutation tests that reflect the actual randomization are design-based analyses for such hypotheses. This means that only such design-based permutation tests can ensure internal validity, without which external validity is irrelevant. However, because of the conservatism of permutation tests, the virtues of permutation tests continue to be debated in the literature, and conclusions are generally of the type that permutation tests should always be used or permutation tests should never be used. A better conclusion might be that there are situations in which permutation tests should be used, and other situations in which permutation tests should not be used. This approach opens the door to broader agreement, but begs the obvious question of when to use permutation tests. We consider this issue from a variety of perspectives, and conclude that permutation tests are ideal to study efficacy in a randomized clinical trial which compares, in a heterogeneous patient population, two or more treatments, each of which may be most effective in some patients, when the primary analysis does not adjust for covariates. We propose the p-value interval as a novel measure of the conservatism of a permutation test that can be defined independently of the significance level. This p-value interval can be used to ensure that the permutation test have both good global power and an acceptable degree of conservatism.

Humans↗

Factor structure and clinical validity of competing models of positive symptoms in schizophrenia.

BACKGROUND: The factor structure of four competing models of positive symptoms and their clinical validity was studied in a sample of 253 schizophrenia inpatients. METHODS: The following models were tested using confirmatory factor analysis: a one-dimension severity model, a two-dimension model comprising a psychosis factor and a disorganization factor, a four-dimension model based on the Scale for the Assessment of Positive Symptoms (SAPS) structure in subscales, and a five-dimension model derived from the previous one by further differentiating Schneiderian delusions from non-Schneiderian ones. RESULTS: More complex multifactorial models fit the data better than simpler models. The five-dimension model was the best adjusted (goodness of fit index = .844, nonnormed fit index = .812, normed fit index = .728). Whereas the one-dimension model did not display significant association with the clinical variables, multidimensional models were related to age at onset and illness severity. The two-dimension model captured well the clinical correlates of the more complex models. CONCLUSION: None of the tested models showed good fit to the data. The one-dimension model displayed both poor factor validity and poor external validity; therefore, research relying on the SAPS total score may reach misleading conclusions.

Adult↗

The future of pharmacoeconomics: bridging science and practice.

In the context of new challenges, issues facing the science, practice, and future of pharmacoeconomics will be discussed. Certain methodologic weaknesses have been observed in published pharmacoeconomic studies, and compromises need to be made between developing an "ideal" method and allowing a study to remain practicable. The objective is to reach a balance between clinical trial-based studies and projective models; trials have high internal validity but low external validity, while models can help explore relevance to real-life settings. Cross-national differences also have an important impact on pharmacoeconomic data; however, using some basic standardized guidelines results from pharmacoeconomic studies may be generalized to other settings. The use of pharmacoeconomic results by decision makers in the United Kingdom has been restrained by unclear priorities within their authority and by the limited availability of credible studies. The future of pharmacoeconomics lies in developing both trial-based and modeling studies, improving their credibility, and meeting the needs of decision makers.

Clinical Trials as Topic↗