PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “External validity”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Validating and updating a prediction rule for neurological sequelae after childhood bacterial meningitis.

Recently, a prediction rule for developing neurological sequelae after childhood bacterial meningitis was developed on a small derivation set. Before implementing in practice a prediction rule must first be tested in new patients (external validation). Our aim was to study the external validity of this rule and, if necessary, to update the rule. The prediction rule was tested on newly available data (validation set) by assessing the rule's calibration and discrimination. We updated the prediction rule by adding extra predictors and re-estimating the regression coefficients of the original predictors in the combined datasets. The rule showed poor agreement between predicted risks and observed frequencies. The ROC area was 0.65 (95% CI 0.57-0.72), which was statistically significantly lower than in the derivation set (0.87 (0.78-0.96)), p-value<0.01. The updated prediction rule showed adequate performance in the combined data sets; the ROC area was 0.77 (95% CI 0.72-0.82). Further study of the generalizability of this updated rule may stimulate application in clinical practice.

Child↗

[Underlying cause of death from external causes: validation of official data in Recife, Pernambuco, Brazil].

OBJECTIVE: To validate the underlying cause of death recorded on the death certificates for individuals under 20 years of age who died from external causes in 1995 in Recife, Pernambuco, Brazil. METHODS: We divided the study into two stages, coding and validation. In both stages we compared the official data concerning causes of death to the data we obtained during our study. We grouped the death certificates into 5 broad categories according to the cause of death; we later subdivided them into 14 categories. We also individually compared the death certificates applying the four-digit system of the International Classification of Diseases, Ninth Revision (ICD-9). We assessed the agreement between the official data and our data in terms of sensitivity and the kappa coefficient. We took as the standard the categorization of the cause of death that we had made during our investigation. RESULTS: In the coding stage, considering all the external causes of death, the overall agreement between the official data and our study data was 94% for the 5 categories, 92% for the 14 categories, and 81% for the four-digit ICD-9 system. In the validation stage the overall agreement was 94% for the 5 categories, 91% for the 14 categories, and 73% for the four-digit ICD-9 system. CONCLUSIONS: Our results suggest that for the death certificates to be reliable, the Institute of Legal Medicine must fill them out following recommended standards. In addition, hospitals and police departments must use greater care in completing the transfer slips that accompany the bodies that are sent to the Institute. More accurate data need to be generated and disseminated for a society to better understand its patterns of violence.

Adolescent↗

Validity of a self-administered food frequency questionnaire (FFQ) and its generalizability to the estimation of dietary folate intake in Japan.

BACKGROUND: In an epidemiological study, it is essential to test the validity of the food frequency questionnaire (FFQ) for its ability to estimate dietary intake. The objectives of our study were to 1) validate a FFQ for estimating folate intake, and to identify the foods that contribute to inter-individual variation of folate intake in the Japanese population. METHODS: Validity of the FFQ was evaluated using 28-day weighed dietary records (DRs) as gold standard in the two groups independently. In the group for which the FFQ was developed, validity was evaluated by Spearman's correlation coefficients (CCs), and linear regression analysis was used to identify foods with large inter-individual variation. The cumulative mean intake of these foods was compared with total intake estimated by the DR. The external validity of the FFQ and intake from foods on the same list were evaluated in the other group to verify generalizability. Subjects were a subsample from the Japan Public Health Center-based prospective Study who volunteered to participate in the FFQ validation study. RESULTS: CCs for the internal validity of the FFQ were 0.49 for men and 0.29 and women, while CCs for external validity were 0.33 for men and 0.42 for women. CCs for cumulative folate intake from 33 foods selected by regression analysis were also applicable to an external population. CONCLUSION: Our FFQ was valid for and generalizable to the estimation of folate intake. Foods identified as predictors of inter-individual variation in folate intake were also generalizable in Japanese populations. The FFQ with 138 foods was valid for the estimation of folate intake, while that with 33 foods might be useful for estimating inter-individual variation and ranking of individual folate intake.

Diet↗

Predicting genotoxicity of aromatic and heteroaromatic amines using electrotopological state indices.

A quantitative structure-activity relationship (QSAR) model relating electrotopological state (E-state) indices and mutagenic potency was previously described by Cash [Mutat. Res. 491 (2001) 31-37] using a data set of 95 aromatic amines published by Debnath et al. [Environ. Mol. Mutagen. 19 (1992) 37-52]. Mutagenic potency was expressed as the number of Salmonella typhimurium TA98 revertants per nmol (LogR). Earlier work on the development of QSARs for the prediction of genotoxicity indicated that numerous methods could be effectively employed to model the same aromatic amines data set, namely, Debnath et al.; Maran et al. [Quant. Struct.-Act. Relat. 18 (1999) 3-10]; Basak et al. [J. Chem. Inf. Comput. Sci. 41 (2001) 671-678]; Gramatica et al. [SAR QSAR Environ. Res. 14 (2003) 237-250]. However, results obtained from external validations of those models revealed that the effective predictivity of the QSARs was well below the potential indicated by internal validation statistics (Debnath et al., Gramatica et al.). The purpose of the current research is to externally validate the model published by Cash using a data set of 29 aromatic amines reported by Glende et al. [Mutat. Res. 498 (2001) 19-37; Mutat. Res. 515 (2002) 15-38] and to further explore the potential utility of using E-state sums for the prediction of mutagenic potency of aromatic amines.

Amines↗

[Interobserver and intraobserver variation: a problem of validity in epidemiologic studies of arterial pressure].

In carrying out blood pressure epidemiologic studies there may be different factors that can affect internal and external validity and thus eliminate the inferential process. As part of the Hypertension and Risk Factors Associated Study conducted in March 1987 in Cuajimalpa de Morelos, Mexico City, 23 nursing students were standardized on the blood pressure auscultatory method using a sound picture and measuring intraobserver and interobserver agreement through intraclass correlation coefficient. Even though initial standardization sessions showed difficulties in the use of instruments and in the reading of blood pressure levels, final K (kappa) values measuring interobserver agreement increased from 0.25 to 0.86. Omega values measuring intraobserver agreement fluctuated between 0.86 and 0.98. This epidemiologic technique is proposed in order to improve internal and external validity of blood pressure studies.

Blood Pressure↗

The WHO (Ten) Well-Being Index: validation in diabetes.

BACKGROUND: In a European trial in 8 countries, the subjective well-being of patients on alternative forms of treatment for insulin-dependent diabetes was compared using the 28-item WHO Well-Being Questionnaire, covering four dimensions of depression, anxiety, energy and positive well-being. The objective of the analysis reported here has been to identify the items of the WHO questionnaire which belong to an overall index of negative and positive well-being. METHODS: Adult patients at 10 study centres in 8 countries who had been on insulin for at least 2 years were invited to participate in a randomised, cross-over trial to compare insulin pump treatment with injection therapy. At each phase, patients completed questions on well-being and general health. Internal validity of the well-being index was evaluated by Cronbach's alpha and Loevinger's and Mokken's homogeneity coefficients, as well as factor analysis. External validity was evaluated by comparisons with results of the general assessment questions and by the ability to discriminate between the alternative forms of treatment. RESULTS: 358 patients had sufficient data for analysis. Ten items were found to constitute a valid index of well-being with respect to internal and external validity. Coefficients of homogeneity were acceptable and there was evidence for both concurrent and discriminant validity. CONCLUSIONS: The WHO (Ten) well-being index includes negative and positive aspects of well-being in a single uni-dimensional scale. Its advantage lies in its ability to show overall change along the continuum of well-being, thus facilitating comparisons between patient groups and treatments. It is not specific to diabetes, and therefore may be useful as a disease-independent index of well-being in a broad range of health care studies.

Adaptation, Psychological↗

Modeling drug albumin binding affinity with e-state topological structure representation.

The binding affinity to human serum albumin for 94 drugs was modeled with topological descriptors of molecular structure, using as experimental data the HPLC chromatographic retention index [logk(HSA)] on immobilized albumin. The electrotopological state (E-State) along with the molecular connectivity chi indices provided the basis for a satisfactory model: r(2) = 0.77, s = 0.29, q(2) = 0.70, s(press) = 0.33. The 10% leave-group-out (LGO) cross-validation method yielded q(2) (= r(2)(press)) = 0.69. Further, the model was tested on a 10 compound external validation set, yielding a mean absolute error, MAE = 0.31; q(2) (= r(2)(press)) = 0.74. MDL QSAR software was used for setting up the data set, creation of combination descriptors, modeling, and database management. All the statistical tests indicate that the topological model is useful for property estimation. Internal and external validation methods were used, and the results indicate that the model is useful for prediction. Randomizations of the activity values also indicate statistically sound models are very different from random statistics. The model indicates that positive factors for binding affinity include electron accessibility and the number of aromatic rings, aliphatic CH groups (-CH(3), -CH(2)-, >CH-), halogens (fluorine and chlorine), and -OH groups. Five-membered heteroatomic rings present a negative factor, whereas six-membered heteroatomic rings present a positive factor. The specific information described can be used as an aid to the drug design process.

Albumins↗

Contemporary identification of patients at high risk of early prostate cancer recurrence after radical retropubic prostatectomy.

OBJECTIVES: To develop a model that will identify a contemporary cohort of patients at high risk of early prostate cancer recurrence (greater than 50% at 36 months) after radical retropubic prostatectomy for clinically localized disease. Data from this model will provide important information for patient selection and the design of prospective randomized trials of adjuvant therapies. METHODS: Proportional hazards regression analysis was applied to two patient cohorts to develop and cross-validate a multifactorial predictive model to identify men with the highest risk of early prostate cancer recurrence. The model and validation cohorts contained 904 and 901 men, respectively, who underwent radical retropubic prostatectomy at Johns Hopkins Hospital. This model was then externally validated using a cohort of patients from the Mayo Clinic. RESULTS: A model for weighted risk of recurrence was developed: R(W)'=lymph node involvement (0/1)x1.43+surgical margin status (0/1)x1.15+modified Gleason score (0 to 4)x0.71+seminal vesicle involvement (0/1)x0.51. Men with an R(W)' greater than 2.84 (9%) demonstrated a 50% biochemical recurrence rate (prostrate-specific antigen level greater than 0.2 ng/mL) at 3 years and thus were placed in the high-risk group. Kaplan-Meier analyses of biochemical recurrence-free survival demonstrated rapid deviation of the curves based on the R(W)'. This model was cross-validated in the second group of patients and performed with similar results. Furthermore, similar trends were apparent when the model was externally validated on patients treated at the Mayo Clinic. CONCLUSIONS: We have developed a multivariate Cox proportional hazards model that successfully stratifies patients on the basis of their risk of early prostate cancer recurrence.

Adult↗

Validation of the Hamilton Depression Rating Scale and Montgommery and Asberg Rating Scales in terms of AGECAT depression cases.

OBJECTIVE: To validate the Hamilton Depression (17) and Montgommery and Asberg Depression Scales as research instruments in older depressed community residents. DESIGN: External validation against GMS/AGECAT case level in the recruitment of older community residents for an antidepressant trial. ANALYSES: Receiver operator curves were generated for each rating scale, using GMS/AGECAT case level in external criterion. The sensitivity, specificity, positive and negative predictive values of both rating instruments were examined in the whole sample and age and gender subgroups. MADRS and HAM-D cut-off scores differentiating GMS/AGECAT cases from subcases were identified. RESULTS: HAM-D cut-off score of 16 and MADRS score of 21 were identified as differentiating case from sub-case. Diagnostic accuracy of both instruments was good, reflecting good sensitivity and specificity across both genders and sub-age groups. CONCLUSIONS: Both scales performed well in this population. These scores provide researchers with externally validated and clinically relevant cut-off scores in designing trials in the management of older depressed community residents.

Aged↗

Lazy structure-activity relationships (lazar) for the prediction of rodent carcinogenicity and Salmonella mutagenicity.

lazar is a new tool for the prediction of toxic properties of chemical structures. It derives predictions for query structures from a database with experimentally determined toxicity data. lazar generates predictions by searching the database for compounds that are similar with respect to a given toxic activity and calculating the prediction from their activities. Apart form the prediction, lazar provides the rationales (structural features and similar compounds) for the prediction and a reliable condence index that indicates, if a query structure falls within the applicability domain of the training database.Leave-one-out (LOO) crossvalidation experiments were carried out for 10 carcinogenicity endpoints ({female/male} {hamster/mouse/rat} carcinogenicity and aggregate endpoints {hamster/mouse/rat} carcinogenicity and rodent carcinogenicity) and Salmonella mutagenicity from the Carcinogenic Potency Database (CPDB). An external validation of Salmonella mutagenicity predictions was performed with a dataset of 3895 structures. Leave-one-out and external validation experiments indicate that Salmonella mutagenicity can be predicted with 85% accuracy for compounds within the applicability domain of the CPDB. The LOO accuracy of lazar predictions of rodent carcinogenicity is 86%, the accuracies for other carcinogenicity endpoints vary between 78 and 95% for structures within the applicability domain.

Algorithms↗

[Clinical usefulness of oligoclonal bands].

The presence of oligoclonal bands (OCB) of immunoglobulin G (IgG) is in our days the most useful finding in the study of the CSF for the diagnosis of multiple sclerosis (MS). The most sensitive method for the detection of OCB is the isoelectric focusing followed by immunoblotting. The prevalence of OCB changes in different populations with a rank of results from 60 to 95 97%. We have determined the prevalence of OCB in our population and the sensitivity and the specificity of the technique used in our laboratory. We have included 391 patients in whom we analysed the presence of OCB, subdivided in; Group 0: Diagnosed of MS, group 1: First episode of demyelinating process, group 2: Neurological disorders considered noninflammatory or nonautoimmune (NINA),group 3: Neurological disorders considered inflammatory, infectious or autoimmune (IIA). The presence of OCB was searched in CSF and serum simultaneously using isoelectric focusing and immunoblotting. In order to standardize the technique we achieved and internal and external validation. Internal validation: sensitivity and specificity (using as a control group first the group NINA and after the group IA). External validation: we choose 10 pairs of CSF/serum from patients with different diagnostics and sent to a reference laboratory ( Karolinska Institute Medical School) that was blind of our results and of the diagnostics. The prevalence of OCB in each group has been: group 0 (MS): 87.7%, group 1: 54.8%, group 2 (NINA): 17.5%, group 3(IIA): 52.7%. Sensitivity: 97.7%, specificity using group NINA as control 82.5% and using group IIA 45.7%. Concordance with the reference laboratory in 9/10 determinations. We conclude that in our population the prevalence of OCB, in patients with MS, is lower than in Northern Europe. The OCB appear in may inflammatory, autoimmune diseases, their specificity for the diagnostic of MS is low.

Autoimmune Diseases↗

Measuring household food security in poor Venezuelan households.

OBJECTIVE: To validate abbreviated methods that estimate food security level among poor communities in Caracas, Venezuela. DESIGN: Two independent cross-sectional studies were undertaken to internally and externally validate simple quantitative/qualitative methods. The quantitative measure was constructed from data on household food availability, gathered using the list-recall method. It is a count of the foods that explain 85% or more of household energy availability. The qualitative measure is a score of female-perceived food insecurity level estimated with a modified 'hunger index', reflecting food resource constraints and hunger experiences within the home. Socio-economic and food behaviour data that may predict household food security (HFS) levels were gathered. The second study was repeated a year later to measure the impact of an increase in the minimum wage on HFS levels. SETTING: Two poor urban communities in Caracas, Venezuela. SUBJECTS: All households in both communities that complied with selection criteria(poor and very poor families that share food resources) and were willing to participate. The sample comprised 238 and 155 female household food managers in the two communities. RESULTS: In 1995, data from females in 238 urban poor households provided evidence for the overall validity of the method. Its application in 1997 to 155 households in the other community gave support to the external validity of the method. Measures were repeated in 1998 on 133 subjects of the above sample, when the minimum wage was increased by 23%. Evidence is presented showing the sensitivity of the method to changes in the determinants of HFS. Data analysed during these three periods suggest that the method can be simplified further by using the food diversity score instead of the quantitative measure since these variables correlate highly with one another(r > or = 2 0.854). CONCLUSIONS: This simple method is a valid and precise measure of food security among poor urban households in Caracas. Th equalitative/quantitative measures complement each other as they capture different dimensions of HFS.

Cross-Sectional Studies↗

Effectiveness, safety and cost-effectiveness of homeopathy in general practice - summarized health technology assessment.

INTRODUCTION: The Health Technology Assessment report on effectiveness, cost-effectiveness and appropriateness of homeopathy was compiled on behalf of the Swiss Federal Office for Public Health (BAG) within the framework of the 'Program of Evaluation of Complementary Medicine (PEK)'. MATERIALS AND METHODS: Databases accessible by Internet were systematically searched, complemented by manual search and contacts with experts, and evaluated according to internal and external validity criteria. RESULTS: Many high-quality investigations of pre-clinical basic research proved homeopathic high-potencies inducing regulative and specific changes in cells or living organisms. 20 of 22 systematic reviews detected at least a trend in favor of homeopathy. In our estimation 5 studies yielded results indicating clear evidence for homeopathic therapy. The evaluation of 29 studies in the domain 'Upper Respiratory Tract Infections/Allergic Reactions' showed a positive overall result in favor of homeopathy. 6 out of 7 controlled studies were at least equivalent to conventional medical interventions. 8 out of 16 placebo-controlled studies were significant in favor of homeopathy. Swiss regulations grant a high degree of safety due to product and training requirements for homeopathic physicians. Applied properly, classical homeopathy has few side-effects and the use of high-potencies is free of toxic effects. A general health-economic statement about homeopathy cannot be made from the available data. CONCLUSION: Taking internal and external validity criteria into account, effectiveness of homeopathy can be supported by clinical evidence and professional and adequate application be regarded as safe. Reliable statements of cost-effectiveness are not available at the moment. External and model validity will have to be taken more strongly into consideration in future studies.

Cost-Benefit Analysis↗

Some methodologic lessons learned from cancer screening research.

Credible and useful methodologic evaluations are essential for increasing the uptake of effective cancer screening tests. In the current article, the authors discuss selected issues that are related to conducting behavior change interventions in cancer screening research and that may assist researchers in better designing future evaluations to increase the credibility and usefulness of such interventions. Selection and measurement of the primary outcome variable (i.e., cancer screening behavior) are discussed in detail. The report also addresses other aspects of study design and execution, including alternatives to the randomized controlled trial, indicators of study quality, and external validity. The authors conclude that the uptake of screening should be the main outcome when evaluating cancer screening strategies; that researchers should agree on definitions and measures of cancer screening behaviors and assess the reliability and validity of these definitions and measures in different populations and settings; and that the development of methods for increasing the external validity of randomized designs and reducing bias in nonrandomized studies is needed.

Biomedical Research↗

Artificial Intelligence for Diagnosis, Risk Stratification, and Prognosis of Neuroblastoma - A Systematic Review and Meta-Analysis.

PURPOSE: To synthesizes evidence on artificial intelligence (AI) performance in neuroblastoma (NB) diagnosis, risk stratification, prognosis, and genomic characterization. MATERIALS AND METHODS: A systematic review and meta-analysis was conducted following PRISMA 2020 guidelines (PROSPERO: CRD42024539475) across five databases. Meta-analyses used random-effects models with logit-transformed Area Under the Curve (AUCs) and cluster-robust standard errors. AI models were classified as Machine Learning Models (MLM) or Hybrid Nomograms (HN) based on their construction methodology. RESULTS: Of 3,742 articles identified, 53 were included. MLMs demonstrated higher point estimates than radiologists in differential diagnosis (AUC: 0.87 vs. 0.83), though this difference was not statistically significant and carried substantial uncertainty. HNs achieved stronger performance in risk stratification (AUC: 0.87). AI-derived nomograms (AUC: 0.9) and gene signatures (AUC: 0.8) outperformed conventional prognostic markers descriptively. Chemotherapy response prediction remained below clinical utility thresholds across all model types. Only 33.9% of models reported calibration and 24.5% underwent external validation. CONCLUSIONS: AI demonstrates proof-of-concept across multiple NB clinical domains. However, clinical adoption remains premature given persistent gaps in external validation, calibration, dataset size, and pediatric-specific model development. Future studies should test these models prospectively in multicenter pediatric cohorts, ideally through COG or SIOPEN, using shared definitions for diagnosis, risk group, treatment response, and survival outcomes.

Humans↗

Popperian epidemiology and the logic of bi-conditional modus tollens arguments for refutational analysis of randomised controlled trials.

Popperian epidemiology is a biomedical science tool based on the hypothesis-deductive method and the falsifiability of scientific hypotheses. This article explores the applicability of the refutationist logic tools in the analysis of a randomised controlled trial (RCT), the randomised Aldactone evaluation study (RALES). This was carried out by using bi-conditional modus-tollens arguments of the type (i) P-then-Q(n) and (ii) Q(n)-If-X(P), X(P) being a set of potential falsifiers of Q(n) as part of the explicit falsity-content of P. In this model, P is the main hypothesis and Q(n) one or more logical predictions to be tested. The X(P) argument represents inclusion criteria, exclusion criteria and conditional criteria of the RCT so every P-then-X(P) argument should be fulfilled in canonical form to corroborate P-then-Q(n). Thus, falsifiability of a RCT would be determined by the empirical content of the conditional argument Q(n)-If-X(P) and its external validity would be determined by the empirical content of X(P). In this way it would be possible to mathematically assess the external validity of a RCT if the observational predicates of the X(P) argument in a given population are known. According to this popperian model, applicability of the RCT results to clinical practice implies transferring of all its empirical content, in other words, the totality of its truth and falsity contents. Thus, to ignore the explicit falsity-content of a RCT such as RALES may jeopardise its potential benefits in clinical practice as suggested by recent studies.

Clinical Trials as Topic↗

Using the Internet to conduct surveys of health professionals: a valid alternative?

OBJECTIVE: The purpose of this study was to examine whether Internet-based surveys of health professionals can provide a valid alternative to traditional survey methods. METHODS: (i) Systematic review of published Internet-based surveys of health professionals focusing on criteria of external validity, specifically sample representativeness and response bias. (ii) Internet-based survey of GPs, exploring attitudes about using an Internet-based decision support system for the management of familial cancer. RESULTS: The systematic review identified 17 Internet-based surveys of health professionals. Whilst most studies sampled from professional e-directories, some studies drew on unknown denominator populations by placing survey questionnaires on open web sites or electronic discussion groups. Twelve studies reported response rates, which ranged from nine to 94%. Sending follow-up reminders resulted in a substantial increase in response rates. In our own survey of GPs, a total of 268 GPs participated (adjusted response rate = 52.4%) after five e-mail reminders. A further 72 GPs responded to a brief telephone survey of non-respondents. Respondents to the Internet survey were more likely to be male and had significantly greater intentions to use Internet-based decision support than non-respondents. CONCLUSIONS: Internet-based surveys provide an attractive alternative to postal and telephone surveys of health professionals, but they raise important technical and methodological issues which should be carefully considered before widespread implementation. The major obstacle is external validity, and specifically how to obtain a representative sample and adequate response rate. Controlled access to a national list of NHSnet e-mail addresses of health professionals could provide a solution.

Attitude of Health Personnel↗

Screening for eating disorders and high-risk behavior: caution.

OBJECTIVE: The current study reviews the state of eating disorder screens. METHODS: Screens were classified by their purported screening function: identification of cases with (a) anorexia nervosa only; (b) bulimia nervosa only; (c) eating disorders in general; (d) partial syndrome, eating disorder not otherwise specified (EDNOS), or subclinical; (e) not a-d but at high risk. Information is presented on development, psychometric properties, and external validation (e.g., sensitivity, specificity, positive predictive values, and negative predictive values). RESULTS: Screens differ widely with regard to objective, psychometric properties and the validation methodology used. Most screens that identify cases are not appropriate for the identification of at-risk behaviors. Little data on the external validity of screens are available. DISCUSSION: Screens should be used with caution. A sequential procedure, in which subjects identified as being at risk during the first stage is followed by more specific diagnostic tests during the second stage, might overcome some of the limitations of the one-stage screening approach.

Feeding and Eating Disorders↗