PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Internal validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Occupational Case Analysis Interview and Rating Scale. An examination of construct validity.

The Occupational Case Analysis Interview and Rating Scale (OCAIRS) was developed based on the Model of Human Occupation with the intention of assessing patients' occupational adaptation. Several studies examining the quality of this instrument have been completed; however, none have discussed the internal validity of the instrument or the appropriateness of the rating scale. The purpose of this study is to validate the internal validity of the OCAIRS and to test the quality of the rating scale. The results indicate that the OCAIRS is a valid measure of occupational adaptation. Each item was shown to have its own rating scale structure, however, all items together still shared the same five-point rating scale.

Adaptation, Psychological↗

[The practice of systematic reviews. III. Evaluation of methodological quality of research studies].

The methodological quality of the primary studies included in a systematic review may influence its results and final conclusions. Methodological quality may be defined in various ways. Partially because of this there are many different assessment lists. The most important dimension of quality is internal validity, defined as the confidence that the design, performance and report of a trial prevent or reduce systematic errors (bias) in the outcomes. For only a limited number of internal validity items a relationship with bias has been proven in empirical studies: concealment of randomisation and blinding of patients and outcome assessors. Preferably, quality should be assessed by at least 2 assessors independently. There is no consensus whether assessment should be done blinded for authors, journal, results and conclusions. Internal validity can be incorporated into statistical pooling in various ways: as a selection criterion, to be used as weight or to hierarchically order studies in a presentation. Well-designed comparative studies are needed to provide clearer guidelines for methodological assessment in the future.

Bias↗

Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies.

PURPOSE: To identify, appraise, and synthesize the performance of artificial intelligence-based meibography reading as compared with human graders in diagnosing meibomian gland dysfunction. METHODS: We followed Cochrane methodology and reporting guidelines for diagnostic test accuracy reviews. To assess potential risk of bias and applicability, we used a modified Quality Assessment of Diagnostic Accuracy Studies-2 checklist. We applied bivariate logistic models to estimate summary sensitivity and specificity when appropriate and used the GRADE framework to rate the certainty of the evidence. RESULTS: We identified 14 eligible studies involving 5511 predominantly middle-aged participants (average age: 27-55 years) who were primarily female (≥54.5%). A total of 18,926 meibography images were obtained through noncontact infrared (11 studies) or in vivo confocal microscopy (three studies). Two studies reported external validation of deep learning models, 12 reported internally validated models, and one reported both. All but one study had high risk of bias in at least one domain; 12 studies raised high or intermediate concern about applicability. Based on three external evaluations, the summary sensitivity and specificity for diagnosing meibomian gland dysfunction from normal glands were 97.5% (95% confidence interval: 77.5%-99.8%) and 85.5% (95% confidence interval: 47.3%-97.5%). Sources of heterogeneity in internally validated models included study population, case mix, and others. The overall evidence was very low to low certainty because of imprecision, high risk of bias, and concerns about applicability. CONCLUSIONS: Artificial intelligence-based meibography grading appears less accurate than human graders. Future studies should adopt rigorous designs, including a more diverse participant pool (or image set), and external validation.

Humans↗

Serial ANCA determinations for monitoring disease activity in patients with ANCA-associated vasculitis: systematic review.

BACKGROUND: Antineutrophil cytoplasmic antibodies (ANCAs) are considered by some investigators to be sensitive markers of disease activity and have been suggested to predict relapse and guide therapeutic decisions. Studies using serial ANCA monitoring in patients with ANCA-associated vasculitis (AASV) have yielded controversial results during the last 15 years. To assess the diagnostic value of serial ANCA testing in the follow-up of patients with AASV, we conducted a systematic review of the available literature. METHODS: Studies were identified by a comprehensive search of the PubMed and BIOSIS+/RRM databases, as well as hand searching. Method quality of all eligible studies was assessed with respect to external and internal validity according to established criteria for diagnostic studies. RESULTS: Twenty-two studies met our inclusion criteria, including a total of 950 patients. Whereas generalizability was not a major problem, assessment of internal validity showed that only a minority of studies reported the combination of consecutive patient recruitment, prospective data collection, and independent determination of both index and reference tests, considered as the ideal for diagnostic test studies. Quantitative meta-analytic calculations were not conducted because of the presence of considerable method heterogeneity. CONCLUSION: The presence of considerable methodological heterogeneity combined with methodological shortcomings with respect to internal validity in the majority of included studies preclude firm conclusions from the available literature concerning the clinical value of serial ANCA determinations for monitoring the follow-up of patients with AASV.

Adult↗

Test-retest reliability of the Holyoake Codependency Index with Australian students.

The Holyoake Codependency Index is a 13-item self-report measure of three aspects of codependency: External Focus, Self-sacrifice, and a sense of being overwhelmed by another person's problematic behavior (termed Reactivity). Previous studies have supported internal validity and the internal consistency and construct validity of the subscales. The present scores for 59 students indicate full scale test-retest reliability of .88 and for subscales (.76 to .82) over a 3-wk. interval.

Adult↗

[Study of the reliability, validity and internal consistency of the LSP scale (Life Skills Profile). Profile of activities of daily living].

The LSP scale (Life Skills profile) has been recently translated and adapted into Spanish language. Its has 39 items and attempts to measure the chronic mental patient functioning in situations and tasks of everyday life. It is brief, jargon free, capable of completion by family members, community housing managers as well as professional staff. Reliability, concurrent validity and internal consistency of this Spanish version are reported. The good performance of LSP in all this measurements give support to its use in research and clinical settings.

Activities of Daily Living↗

Factor structure, concurrent validity, and internal consistency of the Beck Depression Inventory-Second Edition in a sample of college students.

We examined the psychometric properties of the Beck Depression Inventory-Second Edition (BDI-II) [Beck et al., 1996, San Antonio: The Psychological Corporation]. Four hundred fourteen undergraduate students at two public universities participated. A confirmatory factor analysis supported the BDI-II two-factor structure measuring cognitive-affective and somatic depressive symptoms. In addition, the internal consistency was high and the concurrent validity of the BDI-II was supported by positive correlations with self-report measures of depression and anxiety. These findings replicate prior research supporting the validity and reliability of the BDI-II in a college sample.

Adolescent↗

Establishing the internal and external validity of experimental studies.

The information needed to determine the internal and external validity of an experimental study is discussed. Internal validity is the degree to which a study establishes the cause-and-effect relationship between the treatment and the observed outcome. Establishing the internal validity of a study is based on a logical process. For a research report, the logical framework is provided by the report's structure. The methods section describes what procedures were followed to minimize threats to internal validity, the results section reports the relevant data, and the discussion section assesses the influence of bias. Eight threats to internal validity have been defined: history, maturation, testing, instrumentation, regression, selection, experimental mortality, and an interaction of threats. A cognitive map may be used to guide investigators when addressing validity in a research report. The map is based on the premise that information in the report evolves from one section to the next to provide a complete logical description of each internal-validity problem. The map addresses experimental mortality, randomization, blinding, placebo effects, and adherence to the study protocol. Threats to internal validity may be a source of extraneous variance when the findings are not significant. External validity is addressed by delineating inclusion and exclusion criteria, describing subjects in terms of relevant variables, and assessing generalizability. By using a cognitive map, investigators reporting an experimental study can systematically address internal and external validity so that the effects of the treatment are accurately portrayed and generalization of the findings is appropriate.

Double-Blind Method↗

An assessment of calibration and performance of the microdialysis system.

To improve the reliability of microdialysis measurements of tissue concentrations of metabolic substances, this study was designed to test both the performance and the internal validity of the microdialysis methods in the hands of our research group. The stability of the CMA 600 analyser was tested with a known glucose solution in 72 standard microvials and in 48 plastic vials. To evaluate if variation in sampling time makes any difference in sample concentration (recovery), sampling times of 10, 20 and 30 min were compared in vitro with a constant flow rate of 1 microl/min. For testing of sampling times at different flow rates, an in vitro study was performed in which a constant sample volume of 10 microl was obtained. With the no net flux method, the actual concentration of glucose and urea in subcutaneous tissue was measured. The CMA 600 glucose analysis function was accurate and stable with a coefficient of variability (CV) of 0.2-0.55%. There was no difference in recovery for the CMA 60 catheter for glucose when sampling times were varied. Higher flow rates resulted in decreased recovery. Subcutaneous tissue concentrations of glucose and urea were 4.4 mmol/l and 4.1 mmol/l, respectively. To conclude, this work describes an internal validation of our use of the microdialysis system by calibration of vials and catheters. Internal validation is necessary in order to be certain of adequate sampling times, flow rates and sampling volumes. With this in mind, the microdialysis technique is useful and appropriate for in vivo studies on tissue metabolism.

Calibration↗

Quantitation of promoter methylation of multiple genes in urine DNA and bladder cancer detection.

BACKGROUND: The noninvasive identification of bladder tumors may improve disease control and prevent disease progression. Aberrant promoter methylation (i.e., hypermethylation) is a major mechanism for silencing tumor suppressor genes and other cancer-associated genes in many human cancers, including bladder cancer. METHODS: A quantitative fluorogenic real-time polymerase chain reaction (PCR) assay was used to examine primary tumor DNA and urine sediment DNA from 15 patients with bladder cancer and 25 control subjects for promoter hypermethylation of nine genes (APC, ARF, CDH1, GSTP1, MGMT, CDKN2A, RARbeta2, RASSF1A, and TIMP3) to identify potential biomarkers for bladder cancer. We then used these markers to examine urine sediment DNA samples from an additional 160 patients with bladder cancers of various stages and grades and from an additional 69 age-matched control subjects. Data were analyzed on the basis of a prediction model and were internally validated using a jacknife procedure. All statistical tests were two-sided. RESULTS: For all 15 patients with paired DNA samples, the promoter methylation pattern in urine matched that in the primary tumors. Four genes displayed 100% specificity. Of the 175 bladder cancer patients, 121 (69%, 95% confidence interval [CI] = 62% to 76%) displayed promoter methylation in at least one of these genes (CDKN2A, ARF, MGMT, and GSTP1), whereas all control subjects were negative for such methylation (100% specificity, 95% CI = 96% to 100%). A logistic prediction model using the methylation levels of all remaining five genes was developed and internally validated for subjects who were negative on the four-gene panel. This combined, two-stage predictor produced an internally validated ROC curve with an overall sensitivity of 82% (95% CI = 75 % to 87%) and specificity of 96% (95% CI = 90% to 99%). CONCLUSION: Testing a small panel of genes with the quantitative methylation-specific PCR assay in urine sediment DNA is a powerful noninvasive approach for the detection of bladder cancer. Larger independent confirmatory cohorts with longitudinal follow-up will be required in future studies to define the impact of this technology on early detection, prognosis, and disease monitoring before clinical application.

Aged↗

Validation of the pediatric Rome II criteria for functional gastrointestinal disorders using the questionnaire on pediatric gastrointestinal symptoms.

OBJECTIVE: To validate the pediatric Rome II criteria for functional gastrointestinal disorders (FGIDs) using the Questionnaire on Pediatric Gastrointestinal Symptoms (QPGS). METHODS: Subjects were 315 consecutive new patients, 4 to 18 years of age, seen in a tertiary care clinic and classified by pediatric gastroenterologists as having a functional problem. Patients and parents separately completed the QPGS before medical consultation. Diagnoses were derived using computer algorithms reflecting the Rome II criteria for pediatric FGIDs. Convergent validity was assessed by prevalence of diagnoses and internal validity using factor analysis to confirm symptom clusters of the criteria. Separate analyses were performed for 4 to 9 and 10 to 18 year olds, and for diagnoses based on parent and child reports. RESULTS: In both age groups, the most prevalent diagnoses were irritable bowel syndrome (IBS) (22.0%, 35.5%), functional constipation (19.0%, 15.2%), and functional dyspepsia (FD) (13.6%, 10.1%). Parent-child concordance on diagnoses was generally poor. Factor analyses supported the internal validity of FD and of IBS symptoms except for relief with defecation. Although functional abdominal pain syndrome and abdominal migraine occurred rarely, symptom clustering within each diagnosis supports their validity. Among patients with abdominal pain, duration was of at least 3 months in most, and pain was of long duration and severe in at least one third. CONCLUSION: More than half of patients classified as having a functional problem met at least one pediatric Rome II diagnosis for FGIDs. This study offers initial support for the validity of several of the criteria.

Adolescent↗

Development and validation of the Economic Assessment of Glycemic Control and Long-Term Effects of diabetes (EAGLE) model.

BACKGROUND: The Economic Assessment of Glycemic control and Long-term Effects of diabetes (EAGLE) model was developed to provide a flexible and comprehensive tool for the simulation of the long-term effects of diabetes treatment and related costs in type 1 and type 2 diabetes. METHODS: EAGLE simulations are based on risk equations, which were developed using published data from several large studies including the Diabetes Control and Complications Trial, the United Kingdom Prospective Diabetes Study, and the Wisconsin Epidemiological Study of Diabetic Retinopathy. Risk equations for the probability of complications (including hypoglycemia, retinopathy, macular edema, end-stage renal disease, neuropathy, diabetic foot syndrome, myocardial infarction, and stroke) were based on regression analyses, using linear, exponential, and quadratic regression formulae. Subsequent cost calculations are made from the simulated event rates. Internal validation of the EAGLE model was completed by comparing simulated event rates with the published event rates used as the basis for the model. RESULTS: EAGLE provides microsimulations of virtual patient cohorts for type 1 and type 2 diabetes over n years in 1-year cycles. Complications include microvascular and macrovascular events and death, which are calculated over time as cumulative incidences. Glycosylated hemoglobin levels over time are simulated in relation to treatment regimen. Internal validation demonstrated that each mean event rate simulated by EAGLE overlapped with the published mean event (within a range of +/-10%). CONCLUSIONS: The EAGLE model is an evidence-based, internally valid tool for the assessment of the long-term effects of diabetes treatment and related costs.

Diabetes Complications↗

Lessons learned from validation of in vitro toxicity test: from failure to acceptance into regulatory practice.

As no scientific approach or regulatory guidelines existed for the experimental validation of in vitro toxicity tests, in 1990 a US/European validation workshop agreed in Amden (Switzerland) on a simple definition of the validation process. Several international validation studies failed, although they were conducted according to these recommendations. Taking into account the lessons learned from this experience, a second validation workshop was held by ECVAM in Amden in 1994 to develop a more precisely defined validation concept. Prevalidation and the development of biostatistically defined prediction models were added as essential elements to the validation process. In 1995/1996 the ECVAM validation procedure was officially accepted by EU member countries and at the international level by the US regulatory agencies and the OECD. The improved validation concept was immediately introduced into ongoing validation studies. In 1996 the ECVAM/COLIPA validation study of the in vitro phototoxicity test, which was conducted according to the ECVAM/OECD validation concept, was finished successfully and in 1998 a supporting study on UV-filter chemicals was undertaken. In 1998 the 3T3 NRU PT in vitro phototoxicity test was the first experimentally validated in vitro toxicity test that was recommended for regulatory purposes by ESAC, the ECVAM Scientific Advisory Committee, and by the DG ENV of the EU Commission. Meanwhile, two in vitro skin corrosivity tests have successfully been validated by ECVAM. Finally, in June 2000 the three experimentally validated tests were accepted by EU member states for regulatory purposes as the first in vitro toxicity tests. In addition, ECVAM has funded a successful validation study of three in vitro embryotoxicity tests, which was conducted in 12 European laboratories and finished in July 2000. The three tests validated in this study were the whole embryo culture (WEC) test applied to rat embryos, the micromass (MM) test employing primary cultures of dissociated mouse limb bud cells and the mouse embryonic stem cell test (EST). Examples will be given of successful validation studies during the past decade with particular reference to in vitro toxicity tests that were evaluated for regulatory purposes either by the US validation centre ICCVAM or ECVAM in the fields of sensitisation, phototoxicity and embryotoxicity

Animal Testing Alternatives↗

A systematic review of the quality of research on hands-on and distance healing: clinical and laboratory studies.

PURPOSE: To systematically review the quality of published experimental clinical and laboratory research involving hands-on healing and distance healing between 1955 and 2001. DATA SOURCES: Studies were identified through comprehensive literature searches on spiritual healing in MEDLINE, PSYCH LIT, EMBASE, CISCOM, and the Cochrane Library from their inceptions to December 2001. STUDY SELECTION: We selected published randomized, controlled trials of spiritual healing (hands-on healing and distance healing) done in clinical and laboratory settings, all of which had been peer reviewed. DATA EXTRACTION: Independent quality assessment of internal validity was conducted on all identified studies using the comprehensive Likelihood of Validity Evaluation scale. Clinical and laboratory studies were analyzed separately and then subdivided into hands-on healing or distance healing interventions. RESULTS: A total of 45 laboratory and 45 clinical studies published between 1956 and 2001 met the inclusion criteria. Of the clinical studies, 31 (70.5%) reported positive outcomes as did 28 (62%) of the laboratory studies; 4 (9%) of the clinical studies reported negative outcomes as did 15 (33%) of the laboratory studies. The mean percent overall internal validity for clinical studies was 69% (65% for hands-on healing and 75% for distance healing) and for laboratory studies 82% (82% for hands-on healing and 81% for distance healing). Major methodological problems of these studies included adequacy of blinding, dropped data in laboratory studies, reliability of outcome measures, rare use of power estimations and confidence intervals, and lack of independent replication. CONCLUSIONS: When laboratory studies were compared to clinical studies in the areas of hands-on healing and distance healing across the quality criteria for internal validity, distance healing studies scored better than hands-on healing studies, and laboratory studies fared better than clinical studies. Many studies of healing contained major problems that must be addressed in any future research.

Clinical Trials as Topic↗

[Validation of international autoimmune hepatitis group scoring system for diagnosis of type 1 autoimmune hepatitis in Korea].

BACKGROUND/AIMS: There are no pathognomonic features of autoimmune hepatitis (AIH). Its diagnosis requires the exclusion of various other conditions. The aim of this study was to validate indirectly the International Autoimmune Hepatitis Group (IAHG) scoring system in diagnosing AIH. METHODS: Twenty-six patients with Type 1 AIH and female patients with chronic hepatitis B (n=34), chronic hepatitis C (n=25), or toxic hepatitis (n=13) were evaluated according to 9 categories of pretreatment minimum required parameters proposed by IAHG. Aggregate scores of AIH to those of non-AIH groups, which were assessed before and after extracting the proportions of etiologic factors, were also compared and evaluated. RESULTS: While aggregate scores of non-AIH groups, before extracting the proportions of etiologic factors, were 5.2+/-1.8, 5.6+/-1.1, and 7.4+/-1.2 in that order, those of AIH groups were 12.8+/-1.7. These were significantly higher than those of non-AIH groups (p<0.01). All patients in AIH groups and only 1 patient in a non-AIH group showed aggregate scores of more than 10. Aggregate scores after extracting the proportions of etiologic factors were more than 4 in all, except 2, patients. These should have been consistent with 10 if there were no etiologic factors in non-AIH groups. CONCLUSION: The IAHG scoring system might have a relatively excessive importance to the scores of categories excluding distinct etiologies from AIH. It might be difficult to differentiate AIH from chronic liver diseases of indistinct cause based on the IAHG scoring system.

Adult↗

Validation of the French version of the Fecal Incontinence Quality-of-Life (FIQL) scale.

INTRODUCTION: The aim of this multicenter study was to validate the French version of the fecal incontinence quality-of-life scale (FIQL scale) developed in the Unites States of America. PATIENTS AND METHODS: The FIQL scale has 29 items in four scales: lifestyle, coping/behavior, depression/self-perception and embarrassment. Each item is scored from 1 to 4, with poorest quality-of-life scored 1. An average is calculated for each scale. After linguistic validation of the questionnaire, the French version of the FIQL scale was tested twice, at day 0 and day 7, by 100 patients with fecal incontinence (FI). Construction validity, internal reliability, clinical validity and reproducibility were analysed. RESULTS: Analysis of convergent validity of the French version of the FIQL scale showed very good correlation between items and the corresponding scale for lifestyle (0.50-0.79) and depression/self-perception (0.44-0.74), good correlation for coping/behavior (0.31-0.70) and weak correlation for embarrassment (0.30-0.40). Valid discrimination was observed for 24 of the 29 items. Internal reliability was good for each scale (alpha Cronbach between 0.78 and 0.92). Scores determined with the FIQL scale were significantly correlated with Wexner FI scores, demonstrating the clinical validity of the instrument. Reproducibility, evaluated in patients whose FI was unchanged between day 0 and day 7, was good with intraclass correlation coefficients ranging from 0.80 (embarrassment) to 0.93 (lifestyle). CONCLUSIONS: The linguistic and psychometric evaluation demonstrated the validity of the French version of the FIQL scale. This standardized instrument is now available for clinical use in France for quality-of-life assessment in patients with FI.

Adult↗

Prediction model of hepatocarcinogenesis for patients with hepatitis C virus-related cirrhosis. Validation with internal and external cohorts.

BACKGROUND/AIMS: To estimate hepatocarcinogenesis rates in patients with hepatitis C virus (HCV)-related cirrhosis, an accurate prediction table was created. METHODS: A total of 183 patients between 1974 and 1990 were assessed for carcinogenesis rate and risk factors. Predicted carcinogenesis rates were validated using a cohort from the same hospital between 1991 and 2003 (n=302) and an external cohort from Tokyo National Hospital between 1975 and 2002 (n=205). RESULTS: The carcinogenesis rates in the primary cohort were 28.9% at the 5th year and 54.0% at the 10th year. A proportional hazard model identified alpha-fetoprotein (>or=20 ng/ml, hazard ratio 2.30, 95% confidence interval 1.55-3.42), age (>or=55 years, 2.02, 95% CI 1.32-3.08), gender (male, 1.58, 95% CI 1.05-2.38), and platelet count (<100,000 counts/mm3, 1.54, 95% CI 1.04-2.28) as independently associated with carcinogenesis. When carcinogenesis rates were simulated in 16 conditions according to four binary variables, the 5th- and 10th-year rates varied from 9 to 64%, and 21-93%, respectively. Actual carcinogenesis rates in the internal and external validation cohorts were similar to those of the simulated curves. CONCLUSIONS: Simulated carcinogenesis rates were applicable to patients with HCV-related cirrhosis. Since, hepatocarcinogenesis rates markedly varied among patients depending on background features, we should consider stratifying them for cancer screening and cancer prevention programs.

Adult↗

Results of diagnostic accuracy studies are not always validated.

BACKGROUND AND OBJECTIVE: Internal validation of a diagnostic test estimates the degree of random error, using the original data of a diagnostic accuracy study. External validation requires a new study in an independent but similar population. Here we describe whether diagnostic research is validated, which technique is used, and to what extent the validation study results differ from the original. STUDY DESIGN AND SETTING: All original diagnostic accuracy studies published in 1993 in a predefined set of journals were selected. Validation of these studies was assessed in the original article and in articles published within a period of 10 years, through a literature search and contacting the authors. RESULTS: None of the original studies reported any form of validation. Validation studies published later could be identified for 7 of the 11 original studies. Test characteristics were difficult to compare. Despite what was generally believed, not every validation study showed results inferior to the original. We found more studies that evaluated the test in a different population than in a similar one. CONCLUSION: Not every diagnostic accuracy study is validated. Diagnostic tests are more often repeated in different populations.

Databases, Bibliographic↗