PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Internal validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Complex regional pain syndrome: are the IASP diagnostic criteria valid and sufficiently comprehensive?

This is a multisite study examining the internal validity and comprehensiveness of the International Association for the Study of Pain (IASP) diagnostic criteria for Complex Regional Pain Syndrome (CRPS). A standardized sign/symptom checklist was used in patient evaluations to obtain data on CRPS-related signs and symptoms in a series of 123 patients meeting IASP criteria for CRPS. Principal components factor analysis (PCA) was used to detect statistical groupings of signs/symptoms (factors). CRPS signs and symptoms grouped together statistically in a manner somewhat different than in current IASP/CRPS criteria. As in current criteria, a separate pain/sensation criterion was supported. However, unlike in current criteria, PCA indicated that vasomotor symptoms form a factor distinct from a sudomotor/edema factor. Changes in range of motion, motor dysfunction, and trophic changes, which are not included in the IASP criteria, formed a distinct fourth factor. Scores on the pain/sensation factor correlated positively with pain duration (P<0. 001), but there was a negative correlation between the sudomotor/edema factor scores and pain duration (P<0.05). The motor/trophic factor predicted positive responses to sympathetic block (P<0.05). These results suggest that the internal validity of the IASP/CRPS criteria could be improved by separating vasomotor signs/symptoms (e.g. temperature and skin color asymmetry) from those reflecting sudomotor dysfunction (e.g. sweating changes) and edema. Results also indicate motor and trophic changes may be an important and distinct component of CRPS which is not currently incorporated in the IASP criteria. An experimental revision of CRPS diagnostic criteria for research purposes is proposed. Implications for diagnostic sensitivity and specificity are discussed.

Adult↗

The effects of peer review and evidence quality on judge evaluations of psychological science: are judges effective gatekeepers?

Scientifically trained and untrained judges read descriptions of an expert's research in which the peer review status and internal validity were manipulated. Seventeen percent of the judges said they would admit the expert evidence, irrespective of its internal validity. Publication in a peer-reviewed journal also had no effect on judges' decisions. Training interacted with the internal validity manipulation. Scientifically trained judges rated valid evidence more positively than did untrained judges. Untrained judges rated a study with a confound more positively than did trained judges. Training did not affect judge evaluations of studies with a missing control group or potential experimenter bias. Admissibility decisions were correlated with judges' perceptions of the study's validity, jurors' ability to evaluate scientific evidence, and the effectiveness of cross-examination and opposing experts to highlight flaws in scientific methodology.

Adult↗

Predictive validity and internal consistency of the pre-hospital index measured on-site by physicians.

Physiological measures of injury are used as triage tools to identify patients that require treatment in trauma centres. The Pre-Hospital Index (PHI) is based on systolic blood pressure, pulse, respiratory rate, (level of) consciousness, and presence of penetrating injury. The present study evaluated the validity and internal consistency of the PHI. The study was based on 628 patients assessed by physicians at the scene. Mean age was 38.7 years (SD = 24.8), and 65% were male. Motor vehicle collisions caused the injury for 45%. The majority had head/neck (56%) and extremity (45%) injuries. Mean PHI was 4.62 (SD = 5.77), 40% had a PHI of zero, 6% between 1 and 3, 32% between 4 and 7, and 21% greater than 7. The associations between PHI and rates of hospital admission, surgery, ICU treatment, mortality, duration of hospitalization, and length of ICU stay were significant (p < 0.001). A total of 260 (41.4%) patients had major trauma requiring treatment at a trauma centre. A PHI > 3 had 83% sensitivity and 67% specificity for identifying these patients. Internal consistency of the PHI variables was above the acceptable limits. This study has shown that the PHI is a valid and reliable physiological measure of injury severity and field triage tool.

Accidents, Traffic↗

[The practice of systematic reviews. III. Evaluation of methodological quality of research studies].

The methodological quality of the primary studies included in a systematic review may influence its results and final conclusions. Methodological quality may be defined in various ways. Partially because of this there are many different assessment lists. The most important dimension of quality is internal validity, defined as the confidence that the design, performance and report of a trial prevent or reduce systematic errors (bias) in the outcomes. For only a limited number of internal validity items a relationship with bias has been proven in empirical studies: concealment of randomisation and blinding of patients and outcome assessors. Preferably, quality should be assessed by at least 2 assessors independently. There is no consensus whether assessment should be done blinded for authors, journal, results and conclusions. Internal validity can be incorporated into statistical pooling in various ways: as a selection criterion, to be used as weight or to hierarchically order studies in a presentation. Well-designed comparative studies are needed to provide clearer guidelines for methodological assessment in the future.

Bias↗

Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies.

PURPOSE: To identify, appraise, and synthesize the performance of artificial intelligence-based meibography reading as compared with human graders in diagnosing meibomian gland dysfunction. METHODS: We followed Cochrane methodology and reporting guidelines for diagnostic test accuracy reviews. To assess potential risk of bias and applicability, we used a modified Quality Assessment of Diagnostic Accuracy Studies-2 checklist. We applied bivariate logistic models to estimate summary sensitivity and specificity when appropriate and used the GRADE framework to rate the certainty of the evidence. RESULTS: We identified 14 eligible studies involving 5511 predominantly middle-aged participants (average age: 27-55 years) who were primarily female (&#x2265;54.5%). A total of 18,926 meibography images were obtained through noncontact infrared (11 studies) or in vivo confocal microscopy (three studies). Two studies reported external validation of deep learning models, 12 reported internally validated models, and one reported both. All but one study had high risk of bias in at least one domain; 12 studies raised high or intermediate concern about applicability. Based on three external evaluations, the summary sensitivity and specificity for diagnosing meibomian gland dysfunction from normal glands were 97.5% (95% confidence interval: 77.5%-99.8%) and 85.5% (95% confidence interval: 47.3%-97.5%). Sources of heterogeneity in internally validated models included study population, case mix, and others. The overall evidence was very low to low certainty because of imprecision, high risk of bias, and concerns about applicability. CONCLUSIONS: Artificial intelligence-based meibography grading appears less accurate than human graders. Future studies should adopt rigorous designs, including a more diverse participant pool (or image set), and external validation.

Humans↗

[Study of the reliability, validity and internal consistency of the LSP scale (Life Skills Profile). Profile of activities of daily living].

The LSP scale (Life Skills profile) has been recently translated and adapted into Spanish language. Its has 39 items and attempts to measure the chronic mental patient functioning in situations and tasks of everyday life. It is brief, jargon free, capable of completion by family members, community housing managers as well as professional staff. Reliability, concurrent validity and internal consistency of this Spanish version are reported. The good performance of LSP in all this measurements give support to its use in research and clinical settings.

Activities of Daily Living↗

Validity of the 24-hr. dietary recall and seven-day record for group comparisons.

The internal validity of a 24-hr. dietary recall and a seven-day dietary record was investigated among a group of non-institutionalized elderly subjects who were participating in a congregate meals program. Internal validity was assessed by comparing reported intake with unobtrusively obtained data on actual intake. Validity results suggest that the recall is prone to over-reporting low intakes and under-reporting high intakes. This pattern has been referred to as the "flat-slope syndrome." Records collected during the first few dyas were less prone to this syndrome; however, validity declined by the fifth, sixth, and seventh record days. Also, as the record progressed to the seventh day, the demographic nature of the sample became biased due to drop-outs and decreased usability of the records.

Aged↗

The problem of protocol driven costs in pharmacoeconomic analysis.

The increasing number of economic evaluations of healthcare interventions and of drug therapies in particular has been well documented. Surveys of the quality of studies have demonstrated that standards of conduct of such studies have not similarly increased. Concerns over the standards have led to increased calls that economic analyses be more closely linked to randomised controlled clinical trials (RCT). Seven potential threats to the external validity of results limit the generalisability of studies based on RCTs. One such threat is the existence of protocol driven costs. There are two main types of protocol driven costs. Protocol prescribed costs arise as a result of resource use mandated by the clinical trial design. Protocol derived costs occur when increased clinical investigations mandated by trial protocols lead to atypical disease management. Methods to control for protocol driven costs within pharmacoeconomic study designs are available. Modelling studies can be based on data within clinical trials combined with observational data representing more typical resource use. The adoption of pragmatic clinical trial designs provide greater external validity though reduced internal validity. Refinements to explanatory clinical trials can also lead to reduced protocol driven costs. The extent that current studies control for such costs is unclear due to the lack of transparency in the reporting of study methods. A review of published studies found little consideration of protocol driven costs although in several studies there was evidence of their existence. Future studies conducted alongside RCTs should explicitly address how the issue of protocol driven costs was handled within the study framework.

Economics, Pharmaceutical↗

Health care from a behavioral-ecological viewpoint.

While the intended thrust of this paper has been to elicidate the tremendous potential of the behavioral-ecological perspective for health care research and application, the intent has not been to underplay the important role of the biological sciences in the same venture. However, it is my contention that a behavioral-ecological approach to the study of health care has been widely neglected in health care functions and research. In terms of conventional research designs and terminology, the behavioral-ecological research implications can be summarized as follows: a behavioral-ecological perspective of health care research suggests research that is experimental rather than correlational-descriptive; that focuses, because of its naturalistic thrust, on external validity more than internal validity; that incorporates as independent design variables the environmental context in which health behaviors occur; and that allows single-subject as well as multiple-group designs as research strategy. Finally, in terms of dependent variables, the research design requires a clear identification of the observable characteristics of the target health behaviors under consideration. In summary, health research geared toward professional goals appears to profit significantly from an ecological-behavioral approach which provides a model of high explication, specificity, and objectivity for knowledge generation and immediate application.

Behavior Therapy↗

[Quasi experimental evaluation of public health interventions (author's transl)].

The classic experiment, the randomised controlled trial, is the best known and most revered of evaluation research methods. Randomization in community-based intervention trials, however, is not always possible because of ethical problems arising from with holding the experimental treatment from the control groups or the difficulties in conducting experiments in field settings which do not approach controlled laboratory conditions. In such circumstances, quasi-experimental or observational designs must be used. Two major principles are involved in using quasi-experimental methods: (1) the logic for establishing causality between treatment and effect is the same as that for randomised experiments, but the problems of assessing causality or internal validity are greater, and (2) assessment of the external validity or generalizability of quasi-experimental findings crucial to the interpretation of results. Selected quasi-experimental designs using time series and comparison groups are described with examples from public health intervention trials where threats to internal validity have been assessed by using different analytic techniques or gathering additional evidence. Quasi-experimental evaluations are most useful when opportunities exist for testing rival hypotheses concerning the internal and external validity, or the findings can be used to complement true experiments.

Epidemiologic Methods↗

[Validation of French translation of the "Tinnitus Reaction Questionnaire", Wilson et al. 1991].

The present study proposes a validation of a french translation of the TRQ initially published by Wilson et al., for evaluating the psychological distress of tinnitus sufferers. The 26 items translated into french were used on a sample of 173 tinnitus sufferers, who also filled out the Mini-Mult, a short version of the MMPI proposed by Kincannon. Internal validity was demonstrated by strong correlations (i) between each item (except items 5 and 20) and total TRQ score (0.33 < or = r < or = 0.87, p < or = 0.0001), (ii) between each internal TRQ factor (0.58 < r < 0.81, p < 0.0001) and the others. Cronbach's alpha test also showed the questionnaire to have a good internal validity (alpha = 0.94). The external factors used for testing concurrent validity were the scores on depression, psychaesthenia and anxiety Mini-Mult scales. The strong correlations (one factor ANOVA and simple linear regression tests) between scores on depression and psychaesthenia scales and (1) each TRQ item, (2) each TRQ factor, (3) total TRQ scores, confirmed concurrent validity. Scores obtained on anxiety index showed high correlations only with TRQ score, factor 3 score and some TRQ items (most of them included in factor 3). The internal and concurrent validities of the French version of the TRQ justify the use of this questionnaire, with the reserve that items 5 and 20 appeared irrelevant for the measuring of tinnitus distress in French-speaking countries. Such a questionnaire should improve our knowledge of tinnitus' life-impact and enable detection of patients whose psychological distress necessitates rapid intervention.

Adaptation, Psychological↗

Statistical aspects of clinical trials of antibiotics in acute infections.

Controlled clinical trials are important tools for evaluating antibiotics in acute infections. External and internal validity, definition of efficacy criteria, and size of the patient sample constitute special statistical problems in such studies. Critical issues regarding external validity pertain to the selection of patients and to the concept of consecutive patients. The internal validity of a study is influenced by the withdrawal of patients from the evaluation after randomization and the comparability of treatment groups with regard to prognostic factors. The definition of efficacy criteria on the basis of bacteriologic outcomes across control visits is not straightforward. Particularly, the evaluation of efficacy at the last follow-up visit must take into account the accumulated information rather than the cross-sectional information. The most common situation in comparative trials of antibiotics is that rather small differences in efficacy can be anticipated. Sometimes, the question at issue is the demonstration of antibiotic equivalence. For valid conclusions to be made in such situations, large samples must be used. A basic problem affecting many studies of antibiotics is that this criterion is not fulfilled.

Acute Disease↗

Validity of international, time trend, and migrant studies of dietary factors and disease risk.

A linear form relative risk model is used to identify circumstances in which various types of aggregate data lead to valid inferences on relative risk parameters. Upon making a random effects assumption, international or time trend data can lead to appropriate relative risk parameter estimation using iteratively reweighted least-squares procedures. Adequate confounding factor control, however, will typically require data on the distribution of confounding factors in each country or time period. For a simple interpretation of relative risk parameters one may also require data on the joint distribution of primary and confounding factors in each country or time period. Hence disease rate data need to be supplemented by dietary and risk factor survey data in order to avoid confounding bias. Measurement error in individual dietary assessment may, however, limit the ability to quantify the dependence of relative risk on dietary factors, unless the relative risk function is approximately linear in the dietary factors of interest. Most studies of migrant populations involve a comparison of migrant mortality rates with those of their countries of emigration and immigration, with little or no data collection on the dietary habits and risk factors of the migrants themselves. The potential of more comprehensive aggregate data and analytic migrant studies in the diet and disease area is briefly indicated. These issues and methods are illustrated using various types of data pertinent to the association between dietary fat and breast cancer.

Breast Neoplasms↗

Problems translating a questionnaire in a cross-cultural setting.

Questionnaires are widely used in veterinary epidemiological studies. With careful design, assessment and administration, questionnaires can collect accurate data. In cross-cultural settings, questionnaire design is complicated by the added step of translation. As part of a cross-sectional study of smallholder pig raisers in the Philippines, we assessed the importance of translation as a source of error in our questionnaire. The questionnaire contained 120 questions that collected data for entry into 361 separate data-entry fields in a computerised database. The translation was verified using the methods of back-translation and pilot-testing to identify and quantify errors of literal translation, omission and mistranslation. Errors were identified in 44 (37%) questions that obtained data for 147 (41%) data-entry fields. Although errors might have remained in the final translated version of the questionnaire, we have no doubt that the verification procedures substantially improved the internal validity of the questionnaire. It is important that veterinary epidemiologists use such procedures to check the internal validity of translated questionnaires.

Animal Husbandry↗

Testing the effects of nutrient deficiencies on behavioral performance.

The association between specific nutrient deficiencies and poor performance on behavioral tests has been documented for several nutrients. The determination of causality, however, remains elusive. This paper presents the essential criteria for a valid test of causality. Findings from experimental studies in which a nutritional treatment was randomly allocated can be summarized in a statistical statement about the probability that the nutrient treatment caused the behavioral response. Criteria for assessing the internal validity of these studies are examined in terms of whether alleviation of a nutrient deficiency did or did not produce a detectable behavioral response. The plausibility of such a causal inference is dependent on its congruency with known or theorized biological and behavioral mechanisms. External validity describes the extent to which inferences from internally valid studies may be applicable to other populations or circumstances. In addition to these scientific considerations, some of the ethical issues of nutrient-treatment trials are also discussed. All of these considerations provide a better basis for judging whether public health action would be worthwhile than do observed associations that could actually be due to other causes.

Behavior↗

Family planning field research projects: balancing internal against external validity.

This report discusses the experience of a two-year family planning and maternal/child health project in Nepal. Although the project was planned as an experimental field research endeavor, a series of unanticipated events repeatedly compromised the internal validity of the project and forced design changes. While unexpected events are common in the history of most field projects, they present the research evaluator with the fundamental dilemma of trying to maintain a high degree of internal validity without sacrificing external validity. Rigid research designs with tight control over the introduction and measurement of experimental variables may serve to increase internal validity but they may also create an atypical and artificial situation that fails to mirror real field conditions and thus threatens external validity.

Community Health Services↗

Screening for drug abuse among adolescents in clinical and correctional settings using the Problem-Oriented Screening Instrument for Teenagers.

Recent research has indicated high rates of substance abuse among adolescents with emotional and behavioral disorders. Moreover, adolescents in clinical and correctional settings found to have comorbid disorders involving substance abuse experience higher morbidity and mortality rates when compared to adolescents having one or no condition. The present study examines the ability of the Problem-Oriented Screening Instrument for Teenagers (POSIT) to identify DSM-III-R-defined psychoactive substance use disorders among 342 adolescents aged 12-19 years. Participants were sampled from school, clinical, and correctional settings. Optimal-scale cut scores for drug abuse diagnosis classification were derived by a minimum loss function method that minimized false classifications. When using the optimal cut score of two for the total sample, the standard POSIT substance use/abuse scale obtained a drug abuse diagnosis classification accuracy of 84% with sensitivity and specificity ratios of 95% and 79%, respectively. The internal validity of the standard 17-item substance use/abuse scale was subsequently examined by principle component analysis, item analysis, and coefficient alpha. The internal validity analyses were conducted to determine if a shortened scale could be developed and yet retain acceptable classification accuracy. When using the optimal cut score of two for the total sample, the revised 11-item scale obtained a drug abuse diagnosis classification accuracy of 85% with sensitivity and specificity ratios of 91% and 82%, respectively. The results suggest that the POSIT can serve as a useful first-gate instrument to identify adolescents in need of further drug abuse assessment.

Affective Symptoms↗

Internal and external validity in two studies that compared treatment methods.

Research comparing the effectiveness of two treatments offers both strengths and weaknesses for occupational therapy. Although it is worthwhile to determine which of two treatments works best for a particular problem, methodological problems may arise that preclude a valid conclusion. To draw valid conclusions from research, criteria for internal validity and external validity must be satisfied. The two preceding articles in this issue are examples of studies that used between-groups experimental methodology to compare the effectiveness of two different treatments. This paper evaluates the above-mentioned studies on the basis of principles of internal and external validity. One of these studies (Jongbloed, Stacey, & Brighton, 1989) was truly experimental, whereas the other study (Groves & Rider, 1989) was quasi-experimental. Results from both studies were similar because, in each study, both of the treatment groups improved, but there were no significant differences between treatments. Absence of a true control group in both studies presented limitations on the conclusion that both treatments worked equally well.

Analysis of Variance↗