PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “external validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

The Bech-Rafaelsen Mania Scale in clinical trials of therapies for bipolar disorder: a 20-year review of its use as an outcome measure.

Over the last two decades the Bech-Rafaelsen Mania Scale (MAS) has been used extensively in trials that have assessed the efficacy of treatments for bipolar disorder. The extent of its use makes it possible to evaluate the psychometric properties of the scale according to the principles of internal validity, reliability, and external validity. Studies of the internal validity of the MAS have demonstrated that the simple sum of the 11 items of the scale is a sufficient statistic for the assessment of the severity of manic states. Both factor analysis and latent structure analysis (the Rasch analysis) have been used to demonstrate this. The total score of the MAS has been standardised such that scores below 15 indicate hypomania, scores around 20 indicate moderate mania, and scores around 28 indicate severe mania. The inter-observer reliability has been found to be high in a number of studies conducted in various countries. The MAS has shown an acceptable external validity, in terms of both sensitivity and responsiveness. Thus, the MAS was found to be superior to the Clinical Global Impression scale with regard to responsiveness, and sensitivity has been found to be adequate, with the MAS able to demonstrate large drug-placebo differences. Based on pretreatment scores, trials of antimanic therapies can be classified into: (i) ultrashort (1 week) therapy of severe mania; (ii) short-term therapy (3 to 8 weeks) of moderate mania; (iii) short-term therapy of hypomanic or mixed bipolar states; and (iv) long-term (12 months) therapy of bipolar states. The responsiveness of MAS is such that the scale has been able to demonstrated that typical antipsychotics are effective as an ultrashort therapy of severe mania; that lithium and anticonvulsants are effective in the short-term therapy of moderate mania; and that atypical antipsychotics, electroconvulsive therapy (ECT) and transcranial magnetic stimulation seem to have promising effects in the short-term therapy of moderate mania. In contrast, the scale has been used to demonstrate that calcium antagonists (e.g. verapamil) are ineffective in the treatment of mania. MAS has also been used to add to the literature on the evidence-based effect of lithium as a short-term therapy for hypomania or mixed bipolar states and as a long-term therapy of bipolar states.

Antimanic Agents↗

Predicting outcome after traumatic brain injury: development and validation of a prognostic score based on admission characteristics.

The early prediction of outcome after traumatic brain injury (TBI) is important for several purposes, but no prognostic models have yet been developed with proven generalizability across different settings. The objective of this study was to develop and validate prognostic models that use information available at admission to estimate 6-month outcome after severe or moderate TBI. To this end, this study evaluated mortality and unfavorable outcome, that is, death, and vegetative or severe disability on the Glasgow Outcome Scale (GOS), at 6 months post-injury. Prospectively collected data on 2269 patients from two multi-center clinical trials were used to develop prognostic models for each outcome with logistic regression analysis. We included seven predictive characteristics-age, motor score, pupillary reactivity, hypoxia, hypotension, computed tomography classification, and traumatic subarachnoid hemorrhage. The models were validated internally with bootstrapping techniques. External validity was determined in prospectively collected data from two relatively unselected surveys in Europe (n = 796) and in North America (n = 746). We evaluated the discriminative ability, that is, the ability to distinguish patients with different outcomes, with the area under the receiver operating characteristic curve (AUC). Further, we determined calibration, that is, agreement between predicted and observed outcome, with the Hosmer-Lemeshow goodness-of-fit test. The models discriminated well in the development population (AUC 0.78-0.80). External validity was even better (AUC 0.83-0.89). Calibration was less satisfactory, with poor external validity in the North American survey (p < 0.001). Especially, observed risks were higher than predicted for poor prognosis patients. A score chart was derived from the regression models to facilitate clinical application. Relatively simple prognostic models using baseline characteristics can accurately predict 6-month outcome in patients with severe or moderate TBI. The high discriminative ability indicates the potential of this model for classifying patients according to prognostic risk.

Adolescent↗

The feasibility of creating a checklist for the assessment of the methodological quality both of randomised and non-randomised studies of health care interventions.

OBJECTIVE: To test the feasibility of creating a valid and reliable checklist with the following features: appropriate for assessing both randomised and non-randomised studies; provision of both an overall score for study quality and a profile of scores not only for the quality of reporting, internal validity (bias and confounding) and power, but also for external validity. DESIGN: A pilot version was first developed, based on epidemiological principles, reviews, and existing checklists for randomised studies. Face and content validity were assessed by three experienced reviewers and reliability was determined using two raters assessing 10 randomised and 10 non-randomised studies. Using different raters, the checklist was revised and tested for internal consistency (Kuder-Richardson 20), test-retest and inter-rater reliability (Spearman correlation coefficient and sign rank test; kappa statistics), criterion validity, and respondent burden. MAIN RESULTS: The performance of the checklist improved considerably after revision of a pilot version. The Quality Index had high internal consistency (KR-20: 0.89) as did the subscales apart from external validity (KR-20: 0.54). Test-retest (r 0.88) and inter-rater (r 0.75) reliability of the Quality Index were good. Reliability of the subscales varied from good (bias) to poor (external validity). The Quality Index correlated highly with an existing, established instrument for assessing randomised studies (r 0.90). There was little difference between its performance with non-randomised and with randomised studies. Raters took about 20 minutes to assess each paper (range 10 to 45 minutes). CONCLUSIONS: This study has shown that it is feasible to develop a checklist that can be used to assess the methodological quality not only of randomised controlled trials but also non-randomised studies. It has also shown that it is possible to produce a checklist that provides a profile of the paper, alerting reviewers to its particular methodological strengths and weaknesses. Further work is required to improve the checklist and the training of raters in the assessment of external validity.

Bias↗

Validation subset selections for extrapolation oriented QSPAR models.

One of the most important features of QSPAR models is their predictive ability. The predictive ability of QSPAR models should be checked by external validation. In this work we examined three different types of external validation set selection methods for their usefulness in in-silico screening. The usefulness of the selection methods was studied in such a way that: 1) We generated thousands of QSPR models and stored them in 'model banks'. 2) We selected a final top model from the model banks based on three different validation set selection methods. 3) We predicted large data sets, which we called 'chemical universe sets', and calculated the corresponding SEPs. The models were generated from small fractions of the available water solubility data during a GA Variable Subset Selection procedure. The external validation sets were constructed by random selections, uniformly distributed selections or by perimeter-oriented selections. We found that the best performing models on the perimeter-oriented external validation sets usually gave the best validation results when the remaining part of the available data was overwhelmingly large, i.e., when the model had to make a lot of extrapolations. We also compared the top final models obtained from external validation set selection methods in three independent and different sizes of 'chemical universe sets'.

Computer Simulation↗

Performance of a statistical model to predict stroke outcome in the context of a large, simple, randomized, controlled trial of feeding.

BACKGROUND AND PURPOSE: Statistical models to predict the outcome of stroke patients have several uses. Their utility depends on their predictive accuracy in patients other than those on whom they were developed (ie, external validity). We sought to test the external validity of some recently described models in patients enrolled in the FOOD (Feed Or Ordinary Diet) trial: a large randomized trial evaluating feeding policies in patients with stroke. METHODS: The predictive variables were collected during a telephone call to randomize the patient a median of 5 days after stroke onset. Patients were followed up 6 months later to establish their survival, functional status, and residence. Charts were plotted to demonstrate the discrimination and calibration of the models. RESULTS: The models performed well in the first 2955 patients enrolled and followed up in the FOOD trial. The area under the receiver operating characteristic curves varied between 0.78 and 0.81 (with 0.5 indicating no discrimination and 1.0 indicating perfect discrimination). The discrimination was marginally better for patients enrolled within the first day of stroke than later. The models tended to provide rather pessimistic predictions in all groups except those predicted to have a high likelihood of surviving free of dependency. CONCLUSIONS: As one might predict, the discriminatory power in the selected cohort of trial patients was marginally less good than in previously studied unselected cohorts used to test their external validity. These models provide a well-tested tool for stratification in trials, comparing outcomes in different cohorts and examining the additional predictive power of novel factors.

Aged↗

Rapid prediction of acid detergent fiber, neutral detergent fiber, and acid detergent lignin of rice materials by near-infrared spectroscopy.

A rapid predictive method based on near-infrared spectroscopy (NIRS) was developed to measure acid detergent fiber (ADF), neutral detergent fiber (NDF), and acid detergent lignin (ADL) of rice stem materials. A total of 207 samples were divided into two subsets, one subset (approximately 136 samples) for calibration and cross-validation and the other subset for independent external validation to evaluate the calibration equations. Different mathematical treatments were applied to obtain the best calibration and validation results. The highest coefficient of determination for calibration (R2) and coefficient of determination for cross-validation (1-VR) were 0.968 and 0.949 for ADF, 0.846 and 0.812 for NDF, and 0.897 and 0.843 for ADL, respectively. Independent external validation still gave a high coefficient of determination for external validation (r2) and a low standard error of performance (SEP) for the three parameters; the best validation results were SEP = 0.933 and r2 = 0.959 for ADF, SEP = 2.228 and r2 = 0.775 for NDF, and SEP = 0.616 and r2 = 0.847 for ADL, indicating that NIR gave a sufficiently accurate prediction of ADF and ADL content of rice material but a less satisfactory prediction for NDF. This study suggested that routine screening for these forage quality parameters with large numbers of samples is possible with NIRS in early-generation selection in rice-breeding programs.

Calibration↗

The role of artificial intelligence in the diagnosis and prognosis of traumatic brain injury based on brain CT scans: a systematic review.

Traumatic brain injury (TBI) is a leading cause of emergency department visits and a major contributor to injury-related mortality and long-term neurological disability. Non-contrast computed tomography (CT) is the gold-standard imaging modality for the rapid diagnosis of TBI. Clinical outcomes depend strongly on early detection and prompt acute management. Artificial intelligence (AI)-based models may support faster automated identification of traumatic findings and early prediction of patient prognosis.&#xa0;A systematic literature search was conducted in PubMed/MEDLINE, Scopus, IEEE Xplore, ACM Digital Library, and the Cochrane Library in accordance with PRISMA 2020 guidelines to evaluate AI-based models for automated detection of TBI-related findings on CT and for prediction of clinical outcomes. Risk of bias and applicability were assessed using QUADAS-2 for diagnostic accuracy studies and PROBAST&#x2009;+&#x2009;AI for prediction model studies.&#xa0;Twenty-two studies were included. Sixteen studies evaluated diagnostic tasks and 10 evaluated prognostic outcomes, with four studies contributing to both categories. Diagnostic performance was generally high, with many studies reporting AUC values approaching or exceeding 0.90, particularly for larger lesion volumes.Prognostic performance was more variable, with moderate to high discrimination and substantial heterogeneity. Only 9 studies incorporated independent external validation, and performance was frequently lower in external cohorts. All prognostic model studies were judged to be at high overall risk of bias using PROBAST&#x2009;+&#x2009;AI, and most diagnostic accuracy studies also demonstrated high or unclear risk of bias in at least one QUADAS-2 domain, most frequently in patient selection.&#xa0;AI-based models applied to brain CT demonstrate strong technical performance for both diagnostic and prognostic tasks in TBI. However, most studies relied on retrospective designs and lacked independent external validation which limits models generalizability and raises concern for potential overfitting. Prospective, multicenter studies with standardized methodologies and rigorous external validation are required before widespread clinical implementation.

Humans↗

A weakly supervised deep learning-based recurrence prediction and risk stratification of lung adenocarcinoma from pathology whole-slide images.

BACKGROUND: Accurate prediction of postoperative recurrence in lung adenocarcinoma (LUAD) is essential for guiding clinical decision-making and improving patient outcomes. Although various predictive models have been developed, most rely on complex genomic analyses and high-dimensional clinical data. The complexity of these approaches substantially limits their feasibility for routine clinical use. To address this clinical challenge, this study aims to predict postoperative recurrence using routinely available hematoxylin and eosin (H&E)-stained images and characterize the associated biological features. METHODS: A total of 329 patients who underwent curative resection at the First Affiliated Hospital of Wenzhou Medical University (FHWMU) were retrospectively enrolled and randomly assigned to training and internal validation cohorts in a 7:3 ratio. An independent external validation cohort comprising 70 patients from the Clinical Proteomic Tumor Analysis Consortium (CPTAC) was included. Three patch-level feature extractors (Inception_V3, ResNet18, and DenseNet121) were evaluated within a weakly supervised multiple-instance learning (MIL) framework incorporating automated region-of-interest (ROI) detection on segmented whole-slide images (WSIs). Model performance was assessed using the area under the receiver operating characteristic curve (AUC), Kaplan-Meier (KM) survival analysis, and multivariable Cox proportional hazards regression. Transcriptomic profiling and gene set enrichment analysis (GSEA) were conducted to investigate biological differences between risk groups. RESULTS: The model achieved AUCs of 0.923 in the training cohort, 0.891 in the internal validation cohort, and 0.847 in the external validation cohort. The model effectively stratified patients into high- and low-risk groups with significantly different recurrence-free survival (RFS) across all cohorts (all P&#x2009;<&#x2009;0.001) and retained prognostic value within AJCC stages I-III. Transcriptomic analyses revealed consistent enrichment of cell cycle-related pathways and neutrophil extracellular trap (NET) formation in high-risk patients across both institutional and CPTAC cohorts, aligning with distinct biological profiles of the model-derived risk stratification. CONCLUSIONS: This weakly supervised deep learning framework enables accurate and externally validated prediction of postoperative recurrence in LUAD using routinely available histopathological images, and integration of histopathological features with molecular analyses enhances biological interpretability. This work provides a clinically accessible and cost-effective tool for postoperative risk assessment in LUAD patients.

Humans↗

Appropriate methods to assess the effectiveness and efficacy of treatments or interventions to control cancer pain.

Pain is common in cancer patients. To ensure optimal pain management efficacy and effectiveness of new drugs and treatments have to be investigated in clinical trials. Efficacy trials such as randomised controlled trials (RCT) are experimental studies and estimate the maximum potential benefit to be derived from an intervention in ideal circumstances and under a controlled environment. RCTs are the only trial design to establish causal effects. A crossover study is a special type of RCT where patients serve as own controls. In efficacy studies the intervention and the control group should be as homogeneous as possible, confounding variables are controlled, bias is reduced, internal validity is high whereas external validity is low. Studies looking at effectiveness assess clinical practice and reflect real life circumstances. They rely high on external validity at the expense of careful controls, the study population is heterogeneous, confounding variables are examined. Cohort studies follow a group or groups of individuals with a common characteristic over a period of time to measure outcomes. Case-control studies start with the outcome and compare the characteristics of two groups of interest, those with the outcome and those without to identify factors which occur more or less often in the poor outcome group. Definition of outcome criteria is crucial both for efficacy and effectiveness studies and is often a primary problem. All clinical studies must use valid and reliable outcome measures.

Cross-Over Studies↗

A gene-expression signature to predict survival in breast cancer across independent data sets.

Prognostic signatures in breast cancer derived from microarray expression profiling have been reported by two independent groups. These signatures, however, have not been validated in external studies, making clinical application problematic. We performed microarray expression profiling of 135 early-stage tumors, from a cohort representative of the demographics of breast cancer. Using a recently proposed semisupervised method, we identified a prognostic signature of 70 genes that significantly correlated with survival (hazard ratio (HR): 5.97, 95% confidence interval: 3.0-11.9, P = 2.7e-07). In multivariate analysis, the signature performed independently of other standard prognostic classifiers such as the Nottingham Prognostic Index and the 'Adjuvant!' software. Using two different prognostic classification schemes and measures, nearest centroid (HR) and risk ordering (D-index), the 70-gene classifier was also found to be prognostic in two independent external data sets. Overall, the 70-gene set was prognostic in our study and the two external studies which collectively include 715 patients. In contrast, we found that the two previously described prognostic gene sets performed less optimally in external validation. Finally, a common prognostic module of 29 genes that associated with survival in both our cohort and the two external data sets was identified. In spite of these results, further studies that profile larger cohorts using a single microarray platform, will be needed before prospective clinical use of molecular classifiers can be contemplated.

Breast Neoplasms↗

Generalisability of clinical trials in otitis media with effusion.

OBJECTIVE: The impact of a randomised controlled trial (RCT) upon practice depends on its external validity (generalisability). This paper summarises and illustrates a framework for judging and augmenting external validity, emphasising its application to treatment trials in otitis media with effusion (OME) so as to permit stronger inferences in the future. METHODS: The external validity of two surgical trials in the field of OME (TARGET, UK and KNOOP-3, the Netherlands) has been examined within a framework emphasising effect modification, in four specific ways: (1) comparison of the demographic characteristics of the trial population with the domain population; (2) studying the distributions on possible effect modifiers (i.e. variables conditioning the benefit from intervention); (3) studying whether effect modification occurs in the analyses; and (4) comparing outcome measures between the randomised and the eligible but non-randomised children. RESULTS: For neither KNOOP-3 and TARGET were large discrepancies found between randomised and non-randomised children for any of the demographic variables. Differences in distributions along possible effect modifiers were found, but the overlaps were large enough for it still to be possible to study whether these factors indeed modified the outcome. Results for the randomised and non-randomised but eligible patients were similar. The results of both trials therefore appear to be generalisable to their domain populations. CONCLUSIONS: A superficial contrast in the results (KNOOP null, TARGET positive) does not amount to a contradiction, because of differences in the clinical question appropriate to the respective age and populations defined. Attention to quality of design and external validity of randomised controlled trials should achieve higher applicability.

Child↗

Predictive Models for Hypoglycemia Risk in Haemodialysis Patients With Diabetic Kidney Disease: Systematic Review and Meta-Analysis.

AIM: To provide evidence for selecting and developing reliable clinical assessment tools for hypoglycemia in diabetic kidney disease patients during haemodialysis. DESIGN: Review. METHODS: Systematic searches were performed in 9 Chinese and English databases to collect literature regarding the development of hypoglycemia risk prediction models in haemodialysis patients with diabetic kidney disease. Two reviewers independently performed literature screening, data extraction, risk-of-bias assessment, and applicability evaluation. The Prediction Model Risk of Bias Assessment Tool was used to assess the risk of bias and applicability of the included studies. Meta-analysis was conducted using R software. DATA SOURCES: CNKI, Wanfang, VIP, CBM, PubMed, Cochrane Library, EMbase, Web of Science, and CINAHL. The search period covered from the establishment date of each database to December 2025. RESULTS: Six studies, comprising six prediction models, were included. Two studies performed internal validation, and three conducted external validation. All models reported the area under the curve, ranging from 0.813 to 0.866, and calibration measures. Four studies were rated as having a high risk of bias, while all six demonstrated good overall applicability. The meta-analysis showed that the pooled AUC value of the six studies was 0.846 (95% CI: 0.823-0.867). CONCLUSION: Research on hypoglycemia risk prediction models in haemodialysis patients with diabetic kidney disease remains in the developmental stage. Although the included prediction models exhibited satisfactory apparent discriminatory ability and clinical applicability, most of the original studies suffered from a high risk of bias and lacked adequate validation. The true predictive performance and clinical application value of these models remain to be further verified. Accordingly, routine and unconditional clinical application is not recommended at this stage. Future studies should include more high-quality, multicenter external validation and develop models with high generalizability, favourable clinical applicability, and robust predictive performance to facilitate early identification of hypoglycemia risk in this population. IMPACT: This study systematically evaluated the hypoglycemia risk prediction models for diabetic kidney disease patients during haemodialysis, and the research on hypoglycemia risk prediction models for maintenance haemodialysis patients during dialysis is still in the development stage. This study provides a reference for clinical medical staff to select or develop hypoglycemia risk prediction and assessment tools for diabetic kidney disease patients during haemodialysis. REPORTING METHOD: This study was conducted in accordance with the relevant guidelines of the EQUATOR Network and followed the TRIPOD-SRMA Checklist. PATIENT OR PUBLIC CONTRIBUTION: No patient or public contribution. TRIAL REGISTRATION: PROSPERO: CRD420251243352.

Humans↗

Cross-Platform Proteomics and Machine Learning Algorithms Nominate Plasma Biomarkers of Stroke Diagnosis.

BACKGROUND: Blood-based biomarkers for stroke subtyping could improve triage in emergency settings. We used cross-platform proteomics to identify plasma biomarkers differentiating major stroke diagnostic groups. METHODS: We conducted a case-control study using 2 biorepositories. Plasma was collected in the emergency department from adults with suspected stroke before therapeutic intervention. Differentially enriched proteins were identified across acute ischemic stroke, intracerebral hemorrhage, transient ischemic attack, and stroke mimics using SomaScan discovery proteomics (Grady). Differentially enriched proteins were nominated using pairwise and multigroup comparisons and adjusted for clinical covariates. Protein panels were created using least absolute shrinkage and selection operator logistic regression. Internal validation used repeated nested cross-validation (rCV) and targeted mass spectrometry (MS), while external validation used data-independent acquisition &#xa0;mass spectrometry in an independent cohort (Yale). RESULTS: We included 100 subjects (40 with acute ischemic stroke, 20 with intracerebral hemorrhage, 20 with transient ischemic attack, 20 with stroke mimics) in discovery and 80 subjects (20 per group) in external validation cohorts. SomaScan quantified 7307 proteins, of which 61 differentiated stroke subtypes. We identified 7 protein classifiers for acute ischemic stroke (rCV-area under the curve, 0.82 [95% CI, 0.78-0.86]), 6 for intracerebral hemorrhage (rCV-area under the curve, 0.70 [95% CI, 0.64-0.76]), 8 for transient ischemic attack (rCV-area under the curve, 0.78 [95% CI, 0.73-0.84]), and 7 for stroke mimics (rCV-area under the curve, 0.81 [95% CI, 0.77-0.86]). Targeted proteomics internally validated 11 proteins, and data-independent acquisition-mass spectrometry externally validated 32 proteins, including VTN (vitronectin), PLG (plasminogen), and S100A9 as top stroke mimics, transient ischemic attack, and intracerebral hemorrhage classifiers. CONCLUSIONS: This study highlights plasma proteomics as a valuable tool for discovering protein biomarkers of stroke diagnosis. These findings support further validation in larger, multicenter cohorts to facilitate biomarker-guided stroke diagnosis in acute care.

Humans↗

Marital therapy for depression.

BACKGROUND: Marital therapy for depression has the two-fold aim of modifying negative interaction patterns and increasing mutually supportive aspects of couple relationships, thus changing the interpersonal context linked to depression. OBJECTIVES: 1. To conduct a meta-analysis of all intervention studies comparing marital therapy to other psychosocial and pharmacological treatments, or to non-active treatments. 2. To conduct an assessment of the internal validity and external validity. 3. To assess the overall effectiveness of marital therapy as a treatment for depression. 4. To identify mediating variables through which marital therapy is effective in depression treatment. SEARCH STRATEGY: CCDANCTR-Studies was searched on 5-9-2005, Relevant journals and reference lists were checked. SELECTION CRITERIA: Randomised controlled trials examining the effectiveness of marital therapy versus individual psychotherapy, drug therapy or waiting list/no treatment/minimal treatment for depression were included in the review. Quasi-randomised controlled trials were also included. DATA COLLECTION AND ANALYSIS: Data were extracted using a standardised spreadsheet. Where data were not included in published papers, two attempts were made to obtain the data from the authors. Data were synthesised using Review Manager software. Dichotomous data were pooled using the relative risk (RR), and continuous data were pooled using the standardised mean difference (SMD), and 95% confidence intervals (CIs) were calculated. The random effects model was employed for all comparisons. A formal test for heterogeneity, the natural approximate chi-squared test, was also calculated. MAIN RESULTS: Eight studies were included in the review. No significant difference in effect was found between marital therapy and individual psychotherapy, either for the continuous outcome of depressive symptoms, based on six studies: SMD -0.12 (95% CI -0.56 to 0.32), or the dichotomous outcome of proportion of subjects remaining at caseness level, based on three studies: RR 0.84 (95% CI 0.32 to 2.22). In comparison with drug therapy, a lower drop-out rate was found for marital therapy: RR 0.31 (95% CI 0.15 to 0.61), but this result was greatly influenced by a single study. The comparison with no/minimal treatment, showed a large significant effect in favour of marital therapy for depressive symptoms, based on two studies: SMD -1.28 (95% CI -1.85 to -0.72) and a smaller significant effect for persistence of depression, based on one study only. The findings were weakened by methodological problems affecting most studies, such as the small number of cases available for analysis in almost all comparisons, and the significant heterogeneity among studies. AUTHORS' CONCLUSIONS: There is no evidence to suggest that marital therapy is more or less effective than individual psychotherapy or drug therapy in the treatment of depression. Improvement of relations in distressed couples might be expected from marital therapy. Future trials should test whether marital therapy is superior to other interventions for distressed couples with a depressed partner, especially considering the role of potential effect moderators in the improvement of depression.

Depression↗

Validity of residents' self-reported cardiovascular disease prevention activities: the Preventive Medicine Attitudes and Activities Questionnaire.

BACKGROUND: This article describes the development, reliability, and validity of three cardiovascular disease (CVD) prevention subscales-CVD prevention behaviors, perceived importance, and perceived effectiveness-of the Preventive Medicine Attitudes and Activities Questionnaire (PMAAQ). METHODS: The PMAAQ was administered three times to University of Minnesota family practice residents (178) over 2 years (91% response rate). Stability measures were calculated, and validity was demonstrated in four ways: content validity through an expert panel; calculation of internal consistency reliabilities; demonstration of divergent validity; and external validation via a separate chart review. RESULTS: High internal consistency reliabilities among the subscales were seen (Cronbach's alpha = 0.77 to 0.92). Divergent validity was verified by low intercorrelations among the subscales (r = -0.23 to 0.27). Two-month test-retest scores ranged from Cronbach's alpha = 0.47 to 0.64. Significant correlations were seen between the chart review scale and both the CVD behaviors subscale and the PMAAQ smoking scale (r = 0.25 and 0.36, respectively). CONCLUSIONS: Results indicate that the PMAAQ can validly and reliably measure residents' CVD prevention behaviors and provide insight into their preventive health care attitudes. Further, the independence among the subscales suggests that importance and effectiveness by themselves do not affect behavior and that other factors are likely to be important in influencing physician behavior change.

Adult↗

The impact of methodological factors on child psychotherapy outcome research: a meta-analysis for researchers.

Two recent meta-analyses have generated evidence for child and adolescent psychotherapy effects. However, critics note that such meta-analyses often include studies with methodological shortcomings which might invalidate their results. In the present study, we explored whether the results of the most extensive child/adolescent meta-analysis might have been influenced by such methodological variables, focusing on internal validity and external validity factors. Together, these factors accounted for two-thirds as much variance as the substantive factors (e.g., type of therapy, age) in the original meta-analysis. This suggests that relative to these therapy and child-characteristic variables, methodological factors have a substantial, though smaller, impact on meta-analysis results. In general, increased experimental rigor was related to larger effect sizes; this argues against the hypothesis that methodologically weak studies have led to an overestimate of therapy effects. No significant interactive relations were found between validity factors and predictors of outcome; this suggests that the relations noted in previous meta-analyses between outcome and various variables were not distorted by the validity factors tested here.

Adaptation, Psychological↗

Prediction of vapour pressures for halogenated diphenyl ether congeners from molecular descriptors.

BACKGROUND, AIM AND SCOPE: Polychlorinated diphenyl ethers (PCDE) and polybrominated diphenyl ethers (PBDE) have both been identified as environmental contaminants. The physical properties are important in determining the distribution and fate of organic contaminants in the environment. The purpose of the present investigation was to characterise halogenated diphenyl ethers using computationally derived descriptors, and to develop calibration models for the vapour pressure from published experimental data. METHODS: Experimental data for vapour pressures were obtained from the literature. The chemical structure of each PCDE and PBDE congener was optimised prior to descriptor generation. The data analysis was performed using principal component analysis (PCA) and partial least squares regression (PLSR). The calibration models were validated with external test sets. RESULTS AND DISCUSSION: All congeners of PCDEs and PBDEs were characterised by 795 molecular descriptors and two principal components could account for about two thirds of the variance within each group. Bilinear calibration models were developed that could explain 99.4% of the variance in the external validation test sets. Vapour pressures were subsequently predicted for all congeners that were adequately described by these calibration models. The type and number of halogen atoms in the molecule were the main factors influencing the vapour pressures of halogen substituted diphenyl ethers, but the variations in substitution pattern was also shown to be a significant factor. CONCLUSIONS: The molecular descriptor patterns of halogenated aromatic compounds such as diphenyl ethers can be described and interpreted using principal component analysis (PCA). The major sources of variation in the descriptor spaces for PCDEs and PBDEs are the same as those contributing to the differences in vapour pressure, similar to what has previously been reported for the PCBs. The bilinear calibration models for vapour pressure presented here, has a standard error of prediction that is lower than what is reported as the experimental uncertainty or observed as deviations between experimental investigations. The estimated prediction errors are expected to be within the reported boundaries when the models are applied to new objects within the same molecular descriptor space, and model predictions can hence extend the current database of experimental values. RECOMMENDATIONS AND OUTLOOK: The results from this investigation and others show that the establishment of quantitative structure-property relationships (QSPR) is a viable approach to estimate physical properties for halogenated diphenyl ethers. It is easy to foresee an increased need for using QSPR estimation methods in the future, for evaluation of the environmental fate for organic pollutants. Despite method developments and automation, it is unlikely that laboratory determinations can cope with the pace that new pollutants are identified.

Calibration↗

Modeling the toxicity of chemicals to Tetrahymena pyriformis using molecular fragment descriptors and probabilistic neural networks.

The results of an investigation into the use of a probabilistic neural network (PNN)-based methodology to model the 48-h ICG50 (inhibitory concentration for population growth) sublethal toxicity of 825 chemicals to the ciliate Tetrahymena pyriformis are presented. The information fed into the neural networks is solely based on simple molecular descriptors as can be derived from the chemical structure. In contrast to most other toxicological models, the octanol/water partition coefficient is not used as an input parameter, and no rules of thumb or other substance selection criteria are employed. The cross-validation and external validation experiments confirmed excellent recognitive and predictive capabilities of the resulting models and recommend their future use in evaluating the potential of most organic molecules to be toxic to Tetrahymena.

Animals↗