PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “External validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

ADME evaluation in drug discovery. 3. Modeling blood-brain barrier partitioning using simple molecular descriptors.

In this paper, QSPR models were developed for in vivo blood-brain partitioning data (logBB) of a large data set consisting of 115 diverse organic compounds. The best model is based on three descriptors: n-octanol/water partition coefficient calculated using the SLOGP approach, logP; high-charged polar surface areas based on the Gasteiger partial charges, HCPSA, and the excessive molecular weight larger than 360, MW(360). The model bears good statistical significance, n = 78, r = 0.88, q = 0.86, s = 0.36, F = 81.5. The actual prediction potential of the model was validated through two external validation sets of 37 diverse compounds. The predicted results demonstrate that the model bears better prediction potential than many other models and can be used for logBB estimations for drug and drug-like molecules. Comparison of several logP calculation approaches suggests that logP calculated by SLOGP can be used as a significant descriptor for the prediction of molecular transport properties because SLOGP gives the most similar results with CLOGP. The QSPR model indicates that larger polar surface areas have a more negative contribution to logBB, but the absolute partial charges on the atoms surrounded by the polar surfaces should be larger than 0.10|e|. Meanwhile, tight junction membranes limit the size of hydrophilic molecules that can cross the membrane with a molecular weight of approximately 360, because when a molecule's weight is larger than 360 it shows a negative contribution to logBB. The computations of molecular surface, partial charges, logP, and logBB have been accomplished using a program called Drug-BB. Moreover, to improve the efficiency of the computations of logP, we made an extensive reparametrization of SLOGP, and the newly developed SLOGP model is only based on simple atomic addition. Further, we developed a set of parameters to calculate the topological polar surface area (TPSA), thus the high-charged topological polar surface area (HCTPSA) could be estimated from the 2D connection information of a molecule. Adopting the new strategies, the estimations of logP, HCTPSA, and logBB are only based on the topological structure of a molecule and therefore, can be used for fast screening of virtual libraries having millions of molecules.

Absorption↗

ADME evaluation in drug discovery. 5. Correlation of Caco-2 permeation with simple molecular properties.

The correlations between Caco-2 permeability (logPapp) and molecular properties have been investigated. A training set of 77 structurally diverse organic molecules was used to construct significant QSAR models for Caco-2 cell permeation. Cellular permeation was found to depend primarily upon experimental distribution coefficient (logD) at pH = 7.4, high charged polar surface area (HCPSA), and radius of gyration (rgyr). Among these three descriptors, logD may have the largest impact on diffusion through Caco-2 cell because logD shows obvious linear correlation with logPapp (r=0.703) when logD is smaller than 2.0. High polar surface area will be unfavorable to achieve good Caco-2 permeability because higher polar surface area will introduce stronger H-bonding interactions between Caco-2 cells and drugs. The comparison among HCPSA, PSA (polar surface area), and TPSA (topological polar surface area) implies that high-charged atoms may be more important to the interactions between Caco-2 cell and drugs. Besides logD and HCPSA, rgyr is also closely connected with Caco-2 permeabilities. The molecules with larger rgyr are more difficult to cross Caco-2 monolayers than those with smaller rgyr. The descriptors included in the prediction models permit the interpretation in structural terms of the passive permeability process, evidencing the main role of lipholiphicity, H-bonding, and bulk properties. Besides these three molecular descriptors, the influence of other molecular descriptors was also investigated. From the calculated results, it can be found that introducing descriptors concerned with molecular flexibility can improve the linear correlation. The resulting model with four descriptors bears good statistical significance, n = 77, r = 0.82, q = 0.79, s = 0.45, F = 35.7. The actual predictive abilities of the QSAR model were validated through an external validation test set of 23 diverse compounds. The predictions for the tested compounds are as the same accuracy as the compounds of the training set and significantly better than those predicted by using the model reported. The good predictive ability suggests that the proposed model may be a good tool for fast screening of logPapp for compound libraries or large sets of new chemical entities via combinatorial chemistry synthesis.

Caco-2 Cells↗

Prediction of pKa values for aliphatic carboxylic acids and alcohols with empirical atomic charge descriptors.

Two quantitative pKa prediction models for aliphatic carboxylic acids and for alcohols were developed by multiple linear-regression (MLR) analysis with empirical atomic descriptors. The acid and alcohol molecules were described by a set of five and four atomic descriptors, respectively. For the pKa model of 1122 aliphatic carboxylic acids, the squared correlation coefficient is 0.813 with a standard error of prediction of 0.423; for the pKa model of 288 alcohols, the squared correlation coefficient is 0.817 with a standard error of prediction of 0.755, respectively. The good predictive abilities of the models obtained were indicated by both cross-validation and by external validation. An atomic descriptor was developed to model the inductive effect of the neighboring atoms for a central atom in a molecule. The ability of the descriptor to measure the inductive effect of substituent groups was demonstrated by a good correlation of this descriptor with Taft sigma* constants in aliphatic carboxylic acids. It provides a new approach to estimate Taft sigma* constants directly from molecular structures. An algorithm using Kohonen neural networks for splitting a data set into a training set and a test set is also presented.

Alcohols↗

Fast-track failure after cardiac surgery: development of a prediction model.

OBJECTIVE: Risk factors for unsuccessful fast-tracking of cardiac surgery patients have not been collectively defined in the literature. The aim of this study was to determine risk factors for fast-track failure and incorporate them into a predictive fast-track failure score. DESIGN: Prospective observational study. SETTING: Cardiothoracic Department of St Mary's Hospital, London. PATIENTS: Data were collected from April 2003 to April 2005 including 1,084 patients undergoing heart surgery who were admitted into the fast-track unit. INTERVENTIONS: Multifactorial logistic regression was used to develop a propensity score for estimating the likelihood of fast-track failure. MEASUREMENTS AND MAIN RESULTS: One hundred and sixty-nine patients failed fast-track management (15.6%). Independent predictors for fast-track failure were impaired left ventricular function with or without recent acute coronary syndrome (odds ratios 2.89 and 1.65 respectively), re-do operation (one, two, or more vs. none, odds ratio 1.75, 7.98), extracardiac arteriopathy (odds ratio 2.63), preoperative intra-aortic balloon pump (odds ratio 3.09), raised serum creatinine in micromol/L (120-150, >150 vs. <120, odds ratio 1.57, 11.24), and nonelective (odds ratio 3.43) and complex surgery (odds ratio 2.70). Model validation showed very good discrimination (area under the curve = 0.815) and calibration (ĉ statistic = 8.527, p = .129). CONCLUSIONS: The fast-track failure score incorporates several preoperative factors and has been successfully internally validated; after undergoing external validation and possible recalibration it may be used as a tool to facilitate planning and flow of cardiac surgery patients, based on the predicted probability of failure. Application of this score may limit fast-track failure rates and help to reduce morbidity and cost.

Aged↗

Mortality prediction using SAPS II: an update for French intensive care units.

INTRODUCTION: The standardized mortality ratio (SMR) is commonly used for benchmarking intensive care units (ICUs). Available mortality prediction models are outdated and must be adapted to current populations of interest. The objective of this study was to improve the Simplified Acute Physiology Score (SAPS) II for mortality prediction in ICUs, thereby improving SMR estimates. METHOD: A retrospective data base study was conducted in patients hospitalized in 106 French ICUs between 1 January 1998 and 31 December 1999. A total of 77,490 evaluable admissions were split into a training set and a validation set. Calibration and discrimination were determined for the original SAPS II, a customized SAPS II and an expanded SAPS II developed in the training set by adding six admission variables: age, sex, length of pre-ICU hospital stay, patient location before ICU, clinical category and whether drug overdose was present. The training set was used for internal validation and the validation set for external validation. RESULTS: With the original SAPS II calibration was poor, with marked underestimation of observed mortality, whereas discrimination was good (area under the receiver operating characteristic curve 0.858). Customization improved calibration but had poor uniformity of fit; discrimination was unchanged. The expanded SAPS II exhibited good calibration, good uniformity of fit and better discrimination (area under the receiver operating characteristic curve 0.879). The SMR in the validation set was 1.007 (confidence interval 0.985-1.028). Some ICUs had better and others worse performance with the expanded SAPS II than with the customized SAPS II. CONCLUSION: The original SAPS II model did not perform sufficiently well to be useful for benchmarking in France. Customization improved the statistical qualities of the model but gave poor uniformity of fit. Adding simple variables to create an expanded SAPS II model led to better calibration, discrimination and uniformity of fit, producing a tool suitable for benchmarking.

Adult↗

Nomogram for overall survival of patients with progressive metastatic prostate cancer after castration.

PURPOSE: To develop a pretreatment prognostic model for survival of patients with progressive metastatic prostate cancer after castration using parameters that are measured during routine clinical management. PATIENTS AND METHODS: Pretreatment clinical and biochemical determinants from 409 patients enrolled onto 19 consecutive therapeutic protocols from June 1989 through January 2000 were evaluated. The factors selected were age, Karnofsky performance status (KPS), hemoglobin (HGB), prostate-specific antigen (PSA), lactate dehydrogenase (LDH), alkaline phosphatase (ALK), and albumin. These factors were combined in an accelerated failure time regression model to produce a nomogram to predict median, 1-year, and 2-year survival. The nomogram was validated internally and externally using data from a multicenter randomized trial of suramin plus hydrocortisone versus hydrocortisone alone. RESULTS: The median survival of the entire group was 15.8 months (range, 0.9 to 77.8 months); 87% have died. In multivariable analysis, KPS, HGB, ALK, albumin, and LDH were significantly associated with survival (P <.05), whereas age and PSA were not. All seven factors were included in the nomogram. When applied to the external validation data set, the nomogram achieved a concordance index of 0.67. Calibration plots suggested that the nomogram was well calibrated for all predictions. CONCLUSION: A nomogram derived from pretreatment parameters that are measured on a routine basis was constructed. It can be used to predict the median, 1-year, and 2-year survival of patients with progressive castrate metastatic disease with reasonable accuracy. The information is useful to assess prognosis, guide treatment selection, and design clinical trials.

Adult↗

Empirically supported psychosocial interventions for children: an overview.

Discusses issues related to the identification of psychosocial interventions for children that have demonstrated efficacy. Recent debate concerning differences between clinical trials research and clinical practice is summarized, including the tradeoff between interpretability (internal validity) and generalizability (external validity) of outcome studies. This article serves as an introduction to the special issue containing articles that have as their focus the identification of empirically supported psychosocial interventions for children as part of a task force. The article provides an overview of the history, agenda, and methodology used by the task force to define and identify specific empirically supported interventions for children with specific disorders. Whereas a number of well-established or probably efficacious interventions are identified within the series, more work directed at closing the gap between research and practice is needed.

Adolescent↗

Analysis of randomized and nonrandomized patients in clinical trials using the comprehensive cohort follow-up study design.

In clinical research, randomized trials are widely accepted as the definitive method of evaluating the efficacy of therapies. The random assignment of patients to their treatment ensures the internal validity of the comparison of new treatments with controls. An assessment of the external validity of trial results can best be achieved by comparing the study population to the population of patients who met the eligibility criteria but did not consent to randomization. A part of the data of the Coronary Artery Surgery Study (CASS), in which coronary artery bypass surgery is compared to conventional medical therapy in patients with coronary artery disease, is used to illustrate a strategy of multivariate analysis of randomized and nonrandomized patients which allows an investigation of both internal and external validity. The method used Cox's proportional hazards regression model with inclusion of covariates for randomization status and corresponding interactions in addition to the usual covariates for treatment and the important prognostic factors.

Cohort Studies↗

[Validity of the clinical prediction rule for the diagnosis of renal arterial stenosis in hypertensive patients resistant to treatment].

PURPOSE: To perform an external validation of the clinical prediction rule established by Krijnen et al. (Ann Intern Med 1998; 129: 705-11) designed to identify renal artery stenoses (RAS) in hypertensive patients. METHODS: We included 102 patients with a refractory hypertension treated with at least two antihypertensive drugs. All subjects had the research of RAS by renal angiography, or angio-computed tomography, or doppler ultrasound. Probability to detect RAS was calculated with Krijnen's algorithm (Pre-test probability) from the following parameters: age, smoking status, diffuse atherosclerosis, recent hypertension (< 2 y), obesity (BMI > 25), abdominal bruit, hypercholesterolemia (> 6.5 mmol/L), creatinine. ROC curves were plotted for each pre-test probability value. A "post-test probability" was obtained from the likelihood ratio calculated at each pre-test probability level. RESULTS: RAS prevalence in this population was 49%. Area under the ROC curve was 0.79 and Youden index was maximal for a pre-test probability of 15%. Maximal likelihood ratio was obtained for a pre-test probability of 46%. Table shows post-test probability as a function of pre-test probability obtained with Krijnen's algorithm. [table: see text] CONCLUSION: Krijnen's algorithm is valid in a population of resistant hypertensives treated with a bi-therapy. This external validation obtained on a population with a high prevalence of RAS should also be tested on a population with a lower prevalence of SAR.

Age Factors↗

[Theory and practice in medical specialization. I. An instrument for measuring learning strategies].

We present the development and validation of a measurement instrument intended to estimate the degree of vinculation between the theoretical and practical learning activities of medical residents in their usual working conditions in hospitals. The main reason for residents to read medical literature is to find support to their decisions when treating patients. Based on this perspective we designed a self-applied questionnaire that explores diverse circumstances in which the vinculation between theory and practice may be expressed. This instrument was validated through rounds of experts in terms of its construction and content. Its external validity was explored with two groups of internal medicine residents with a different degree of vinculation between theory and practice: one high and the other with a low vinculation. In addition, we designed a guide for the direct observation of the theoretical and practical clinical learning activities in order to estimate its concordance with the results of the questionnaire. The questionnaire was able to discriminate the group differences and showed a satisfactory concordance with the information provided by direct observation. A copy of the questionnaire is available by request to the authors. We conclude that in its present stage of development, the instrument has shown internal and external validity and may be used to explore the process training of medical residents.

Education, Medical↗

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning↗

Inhibition of the tyrosine kinase, Syk, analyzed by stepwise nonparametric regression.

A set of 538 inhibitors of the tyrosine kinase, Syk, including purines, pyrimidines, indoles, imidazoles, pyrazoles, and quinazolines, has been analyzed using a stepwise nonparametric regression (SNPR) algorithm, which has been developed for QSAR studies of pharmacological data. The algorithm couples stepwise descriptor selection with flexible, nonparametric, kernel regression, to generate structure-activity relationships. A further 371 molecules have been used as a test set to evaluate the models generated. Descriptors were selected using an internal monitoring set, and models were assessed using 10% of the principal (538-compound) data set, selected randomly, as an external validation set. The best model had a Q(2) of 0.46 for the external validation set. Test set predictions were significantly less accurate, partly due to the higher mean activity of the test molecules. However at a more coarse-grain level the SNPR models classified active molecules accurately, giving good enrichments. The data sets are difficult to model accurately and SNPR performs better than multilinear regression and a neural network analysis. In the additive implementation of SNPR multidimensional models are considered as a sum of single dimensional regressions. This makes the resultant models easily interpretable. For example, in the most predictive SNPR models, there is a clear nonlinear relationship between hydrophobicity (AlogP98) and inhibitory activity.

Enzyme Inhibitors↗

Screening for controlled substance abuse in interventional pain management settings: evaluation of an assessment tool.

There is a need for an assessment tool to identify drug abuse behaviors in patients in pain treatment practices. Many assessment tools are complex, lengthy, lack external validation, and/or are difficult to administer. This prospective evaluation was undertaken to provide external validation for an assessment tool with 12 sections and 27 items. The test was applied in a prospective fashion to 500 consecutive patients: 100 patients in a drug abuse group and 400 patients in a non-abuse group. Drug abuse was defined as the misuse of controlled substances in a clinical setting, including obtaining controlled substances from other physicians or other identifiable sources, dose escalation with inappropriate use, and/or violation of controlled substance agreements. This study was performed in an interventional pain management setting with patients who were in stable therapy and were followed for at least one year. Results identified 8 of 12 parameters to be useful in identifying patients with drug abuse. Three factors were particularly useful, allowing correct identification of patients with abuse behavior in 90% of cases (odds ratios greater than 100 and P values of 0.001 or less). Important factors identified included excessive opiate needs, deception or lying to obtain controlled substances, and current or prior intentional doctor shopping. Together, these factors appear to identify 90% of patients with drug abuse. This tool provides a simple, reliable, and cost effective means of screening for drug abuse during the clinical evaluation of patients in interventional pain management settings.

Journal Article↗

Representativeness and response rates from the Domestic/International Gastroenterology Surveillance Study (DIGEST).

BACKGROUND: The Domestic/international Gastroenterology Surveillance Study (DIGEST) examined the prevalence of upper gastrointestinal symptoms among the general population in 10 countries, and the impact of these symptoms on healthcare usage and quality of life. This report discusses the validation of the DIGEST sample and reviews the response rates from the survey. METHODS: External validation of the DIGEST sample was conducted by comparing the age, age by gender and annual household incomes of the sample with census-derived data. A comparison was also made between Psychological General Well-Being Index (PGWBI) scores from study subjects in the Scandinavian countries and the USA and the total sample population norms. RESULTS: Under- and oversampling, defined as > or =5% difference from the population norms, was evident in eight out of 10 countries, but no systematic bias was evident. The final distribution of the sample by gender was 51% female and 49% male. Although differences in PGWBI scores were noted between DIGEST subjects and population norms, these differences were <0.30 standard deviations--markedly below the difference considered as relevant for the PGWBI. Response for the survey in individual countries ranged from 17% in the USA to 61% in Norway, with a survey-wide rate of 27%. The overall response rate, including primary non-respondents, was 13.4%. The majority of nonresponse (51.4%) was attributed to failure to establish contact with the subjects, with 41.7% of subjects declining to be interviewed and the remaining 6.9% of subjects not meeting the age and sex criteria used for the survey. CONCLUSIONS: The DIGEST sample exhibited good external validity, providing a foundation for comparison between data derived from individual countries in the survey.

Adult↗

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n&#xa0;=&#xa0;907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n&#xa0;=&#xa0;35), colorectal cancer (n&#xa0;=&#xa0;21), and pancreatic cancer (n&#xa0;=&#xa0;9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans↗

How generalizable are the effects of smoking prevention programs? Refusal skills training and parent messages in a teacher-administered program.

This study investigated both substantive and methodological issues associated with school-based smoking prevention programs. Substantive issues included the efficacy of a refusal skills training curriculum and of parent messages mailed to students' homes. Methodological issues included the effects of assigning classrooms versus entire schools to experimental conditions and determination of the effects of attrition on internal and external validity. Results revealed differential impact for different subgroups of adolescents. The refusal skills program produced lower rates of smoking than the control condition for students who were smokers at the pretreatment assessment but may have produced detrimental effects among males who were nonsmokers at pretest. The provision of parent messages did not affect outcome. Method of assignment (schools versus classrooms) failed to produce significant effects, and attrition did not affect internal validity. However, the above differential findings, as well as the impact of attrition on external validity, raise questions concerning the generalizability of smoking prevention programs.

Adolescent↗

A disease activity score for polymyalgia rheumatica.

OBJECTIVE: To develop a composite score for measurement of disease activity in polymyalgia rheumatica (PMR) and assess its internal and external validity. METHODS: A PMR activity score (AS) was designed and assessed for internal and external validity in two patient cohorts: 57 international patients evaluated primarily for development of the PMR-AS at baseline, weeks 4 and 24; and for validation, 24 Austrian patients assessed at baseline, week 4, and at a mean (SD) point of week 33.6 (24.5). The PMR-AS was calculated as: CRP (mg/dl)+VAS p (0-10)+VAS ph (0-10)+(MST (min)x0.1)+EUL (3-0); Cronbach's alpha was calculated. Factor analysis by linear regression was applied, and responses calculated on the basis of the PMR response criteria and the PMR-AS applied. PMR-AS values at different times were compared by paired t tests. RESULTS: Cronbach's alpha for the composite score was 0.91 and 0.88 in the two cohorts. Factor analysis showed that each single item contributed significantly to the total score and the relative weight of each item in both cohorts was equally distributed. Mean PMR-AS at baseline was 27.54 and 28.72, respectively, at week 4, 5.99 and 8.99, and at the final visit 5.35 and 5.92 (NS). PMR-AS values at baseline and at later visits were significantly different (p<0.0001). PMR-AS values <7 indicated low disease activity, 7-17 medium disease activity, and >17 high PMR activity. In a third control cohort the PMR-AS correlated highly with patient's global assessment, patient satisfaction, and ESR (p<0.001). CONCLUSION: The PMR-AS provides an easily applicable and valid tool for monitoring disease activity, and in combination with the PMR response criteria provides a better description of response.

Aged↗

Design issues for conducting cost-effectiveness analyses alongside clinical trials.

In response to rising demands for timely economic data on new medical technologies, cost-effectiveness studies are increasingly being conducted alongside clinical trials. Because of the historical differences in perspective and methods between cost-effectiveness studies and clinical trials, the design phase of these hybrid trials requires special consideration. Cost-effectiveness studies require more comprehensive evaluations of outcomes than the endpoints typically measured in clinical trials. Often, these comprehensive outcome measures (such as quality of life) prove useful for interpreting the other endpoints measured in the trial, as well as for estimating the cost-effectiveness of the intervention. In this manuscript, we discuss several aspects related to the design of joint clinical/economic trials, including study perspective, hypothesis testing, sample size estimation, and methods for collecting cost and outcome data. We also discuss issues that may limit the external validity of the cost-effectiveness results of these trials. Many potential threats to external validity can be successfully addressed if they are identified and accounted for in the design phase of the study.

Clinical Trials as Topic↗