PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Internal validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Prognostic nomogram for renal insufficiency after radical or partial nephrectomy.

PURPOSE: We analyzed prognostic factors to predict renal insufficiency after partial or radical nephrectomy. We developed and performed internal validations of a postoperative nomogram for this purpose. We used a prospectively updated renal tumor database of more than 1,500 patients. MATERIALS AND METHODS: From July 1989 to October 2003, 161 partial nephrectomies and 857 radical nephrectomies performed at Memorial Sloan-Kettering Cancer Center for renal cortical tumors were analyzed. Computerized tomography images were reviewed by a single radiologist. Kidney volume was calculated using the ellipsoid formula, V = L1 x L2 x L3 x pi/6, where V represents volume and L represents length. Renal insufficiency was defined by 2 serum creatinine values greater than 2.0 mg/dl at least 1 month postoperatively. Tumor histology was not an exclusion criterion and yet we excluded cases of bilateral synchronous disease. Prognostic variables were preoperative serum creatinine, American Society of Anesthesiologists score, percent change in kidney volume after surgery, and patient age and sex. RESULTS: Renal insufficiency was noted in 105 of the 857 patients with radical nephrectomy (12.3%) and in 6 of the 161 with partial nephrectomy (3.7%) studied. Patients had a median followup of 21.2 months (maximum 157.9). The 7-year probability of freedom from renal insufficiency in the cohort was 79.1% (95% CI 74.6 to 83.6). The nomogram was designed based on a Cox proportional hazards regression model. Following internal statistical validation nomogram predictions appeared accurate and discriminating with a concordance index of 0.835. CONCLUSIONS: A nomogram was developed that can predict the 7-year probability of renal insufficiency in patients undergoing radical or partial nephrectomy.

Adolescent↗

Making sense of ambiguity: evaluation in internal reliability and face validity of the SF 36 questionnaire in women presenting with menorrhagia.

OBJECTIVE: To determine the face validity and internal reliability of the short form 36 (SF 36) health survey questionnaire in women presenting with menorrhagia. DESIGN: Postal survey of women recruited by their general practitioners followed by interviews of a selected subsample. PATIENTS: 348 women who had consulted their general practitioner with excessive menstrual bleeding and completed questionnaires after treatment. 49 women selected from this group were interviewed in depth about their health status, and requested to complete the SF 36 questionnaire. MAIN MEASURES: Subjective accounts of functioning and wellbeing as measured by the eight scales of the SF 36 questionnaire. RESULTS: Data from the postal survey indicated that the ¿general health perceptions¿ and ¿mental health¿ scales of the SF 36 questionnaire had lower internal reliability coefficients than documented elsewhere. In the follow up interviews several questions on the SF 36 questionnaire were commented on as inappropriate or difficult to answer for patients with heavy menstrual bleeding. CONCLUSIONS: Some questions on the SF 36 questionnaire were difficult to answer for this group of patients. Such problems can adversely effect the validity of the measure. It is suggested that comments of patients upon measures such as the SF 36 questionnaire could both determine the appropriateness of such measures for given studies and influence questionnaire design.

Adult↗

Cost-efficient study designs for binary response data with Gaussian covariate measurement error.

When mismeasurement of the exposure variable is anticipated, epidemiologic cohort studies may be augmented to include a validation study, where a small sample of data relating the imperfect exposure measurement method to the better method is collected. Optimal study designs (i.e., least expensive subject to specified power constraints) are developed that give the overall sample size and proportion of the overall sample size allocated to the validation study. If better exposure measurements can be collected on a sample of subjects, an optimal design can be suggested that conforms to realistic budgetary constraints. The properties of three designs--those that include an internal validation study, those where the validated subsample is derived from subjects external to the primary investigation, and those that use the better method of exposure assessment on all subjects--are compared. The proportion of overall study resources allocated to the validation substudy increases with increasing sample disease frequency, decreasing unit cost of the superior exposure measurement relative to the imperfect one, increasing unit cost of outcome ascertainment, increasing distance between two alternative values of the relative risk between which the study is designed to discriminate, and increasing magnitude of hypothesized values. This proportion also depends in a nonlinear fashion on the severity of measurement error, and when the validation study is internal, measurement error reaches a point after which the optimal design is the smaller, fully validated one.

Cohort Studies↗

Validation studies and proficiency testing.

Genetically modified organisms (GMOs) entered the European food market in 1996. Current legislation demands the labeling of food products if they contain <1% GMO, as assessed for each ingredient of the product. To create confidence in the testing methods and to complement enforcement requirements, there is an urgent need for internationally validated methods, which could serve as reference methods. To date, several methods have been submitted to validation trials at an international level; approaches now exist that can be used in different circumstances and for different food matrixes. Moreover, the requirement for the formal validation of methods is clearly accepted; several national and international bodies are active in organizing studies. Further validation studies, especially on the quantitative polymerase chain reaction methods, need to be performed to cover the rising demand for new extraction methods and other background matrixes, as well as for novel GMO constructs.

Calibration↗

Preoperative erectile function is one predictor for post prostatectomy incontinence.

AIMS: The precise etiology of post prostatectomy incontinence (PPI) is not fully understood and risk factors are not yet comprehensively defined. It has been reported that sparing of the neurovascular bundle during prostatectomy improves postoperative erectile function, whereas the influence on urinary control is unclear. From daily clinical experience we made the impression that patients who are in the best shape have better erections and better continence. We therefore searched our database for a possible correlation between the preoperative erectile function and the incidence of PPI. PATIENTS AND METHODS: Four hundred three patients who underwent radical retropubic prostatectomy between January 2000 and May 2003 were enrolled into this retrospective study. Data of 327 patients (response rate 81%) at a median follow-up of 26 months were analyzed using the validated International Index of Erectile Function (IIEF 5), the validated Urinary Distress Inventory (UDI6) and a standardized urinary symptom inventory. Continence was defined as usage of no or one pad daily. Erectile Dysfunction (ED) was defined as none/mild or moderate/severe with an IIEF 5 score of 17 or more or less than 17, respectively. RESULTS: Univariate and mulitvariate logistic regression analysis including preoperative IIEF 5 scores, age and nerve sparing prostatectomy, identified preoperative erectile function as significant predictor for PPI (P = 0.024), whereas age (P = 0.759) and nerve sparing prostatectomy (P = 0.504) did not predict PPI. CONCLUSION: Erectile function is a predictor of PPI and should be recorded preoperatively.

Aged↗

The BREV neuropsychological test: Part I. Results from 500 normally developing children.

The Battery for Rapid Evaluation of Cognitive Functions (Batterie Rapide d'Evaluation des Fonctions Cognitives: BREV) was designed to provide health professionals with a quick clinical tool for screening acquired and developmental cognitive deficits in children aged 4 to 8 years. The BREV explores oral language in both its expressive and receptive forms, non-verbal functions, attention, verbal and visuo-spatial memory, and main learning acquisition. Results of the first phase of validation are presented in this report consisting of internal validity measurements gained by testing 500 normally developing school children (257 females, 243 males; mean age 6 years 7 months, SD 1 year 6 months. The validation provides appropriate values for each of the 17 subtests assessing cognitive functions (oral language, non-verbal abilities, attention and memory, educational achievement) in 10 age groups, from 4 to 8 years of age. All subtests with the same content for any age revealed values which increased significantly with age. Interreliability was tested in a retest for 70 children and scores obtained on retesting correlated significantly with initial values. The BREV is a reliable test with carefully established normative values, appropriate for preschool and school-age children.

Attention↗

A production task evaluation of individual differences in mental addition skill development: internal and external validation of chronometric models.

A production task paradigm for obtaining reaction times to mental addition stimuli was used for internal and external validation of chronometric models of mental addition processing. The first analysis explored the internal validity of extant chronometric models and found that three models, (a) a tabular memory network retrieval strategy (PRODUCT), (b) a nontabular memory network retrieval strategy (ERROR RATE), and (c) a computational strategy (MIN), were able to encompass individual differences in strategy choice for 155 individuals from Grades 2 to 8 and 111 college students. Patterns of convergent and discriminant validity for these models were also demonstrated. The second analysis explored the external validity of relations among (a) two traditionally measured factor analytic dimensions of ability, Numerical Facility and Perceptual Speed; (b) two information processing dimensions presumed to underlie mental addition. Addition Efficiency and Speediness; and (c) a digit-span measure of Short-Term Memory. We specified a series of two-group (grade school and college) structural equation models to represent the relations among all measures and showed that individual differences in the apparently calculative processes that underlie the traditionally defined ability dimension of Numerical Facility are highly related to individual differences in Addition Efficiency and Speediness of information processing.

Adolescent↗

The WHO (Ten) Well-Being Index: validation in diabetes.

BACKGROUND: In a European trial in 8 countries, the subjective well-being of patients on alternative forms of treatment for insulin-dependent diabetes was compared using the 28-item WHO Well-Being Questionnaire, covering four dimensions of depression, anxiety, energy and positive well-being. The objective of the analysis reported here has been to identify the items of the WHO questionnaire which belong to an overall index of negative and positive well-being. METHODS: Adult patients at 10 study centres in 8 countries who had been on insulin for at least 2 years were invited to participate in a randomised, cross-over trial to compare insulin pump treatment with injection therapy. At each phase, patients completed questions on well-being and general health. Internal validity of the well-being index was evaluated by Cronbach's alpha and Loevinger's and Mokken's homogeneity coefficients, as well as factor analysis. External validity was evaluated by comparisons with results of the general assessment questions and by the ability to discriminate between the alternative forms of treatment. RESULTS: 358 patients had sufficient data for analysis. Ten items were found to constitute a valid index of well-being with respect to internal and external validity. Coefficients of homogeneity were acceptable and there was evidence for both concurrent and discriminant validity. CONCLUSIONS: The WHO (Ten) well-being index includes negative and positive aspects of well-being in a single uni-dimensional scale. Its advantage lies in its ability to show overall change along the continuum of well-being, thus facilitating comparisons between patient groups and treatments. It is not specific to diabetes, and therefore may be useful as a disease-independent index of well-being in a broad range of health care studies.

Adaptation, Psychological↗

Scholarly literature review: Efficacy of psychological interventions for pediatric chronic illnesses.

OBJECTIVE: To review empirical studies of the efficacy of psychological interventions as adjuvant therapies for children with pediatric diabetes, cancer, cystic fibrosis, and sickle cell disease. METHODS: A search was conducted for qualifying studies published since 1980. Only studies meeting basic criteria for external and internal validity were included. Nineteen studies were identified, providing data on 62 outcome variables. Effect sizes (ESs) were analyzed by illness type, intervention type, and strength of internal and external validity of the research design. RESULTS: Overall, interventions were associated with large ESs, which were not significantly moderated by illness type or intervention type. However, larger ESs were associated with lower scores on validity of research design. CONCLUSIONS: Adjuvant psychological interventions for pediatric chronic illnesses appear in general to be efficacious, associated with a large mean ES across a range of outcome variables. However, until more studies have been completed using stronger research designs, only tentative conclusions can be drawn.

Adaptation, Psychological↗

Randomized database studies: a new method to assess drugs' effectiveness?

The need to evaluate drugs' effects in real clinical practice is increasingly important. Randomized clinical trials (RCTs) and database analyses (DBA) are the two main methods to assess treatments effectiveness. RCTs remain the "gold standard" for comparing alternative treatments. However, they are conducted under strict, protocol-driven conditions that may limit their generalizability. Advantages of new high quality clinical databases, on the other hand, include the simple and economic access to large number and range of cases, and the ability to capture all aspects of actual medical practice. The main potential limitation of DBA is the potential for comparison bias due to the lack of randomization. Despite the efforts to design naturalistic trials and to use sophisticated statistical techniques to minimize selection bias, the inherent limitations of both methods (problems of external and internal validity, respectively) have not been completely solved. Thus, the actual challenge is the development of some new strategy capable of generating results with an acceptable balance between internal and external validity. As randomization is essential to minimize comparison bias, we point out the possibility to include randomization modules in computer-based patient records. The theoretical foundation of these "randomized database studies" is the simultaneous use of both experimental and observational methods in the assessment of drugs' effectiveness. The progressive standardization of clinical practice and the development and adoption of improved computer-based patient records could facilitate the use of this new research strategy.

Databases, Factual↗

Menopausal Vasomotor Symptoms (MVS) survey for assessment of hot flashes.

OBJECTIVES: During the menopause, 65%-80% of women experience hot flashes. Hot flashes can also occur after hysterectomy and may be experienced during cancer therapy. A very limited assessment of hot flashes is provided by currently available menopausal instruments. This study was performed to evaluate a new Menopausal Vasomotor Symptoms (MVS) survey as an instrument for a comprehensive and subjective assessment of hot flashes. METHODS: The MVS survey was designed to assess multiple dimensions of hot flashes and to be simple and easy to administer. Hot flashes and associated conditions were addressed with 39 closed-ended questions. Sixty-one qualified women, 40-58 years old, from the Dallas, Fort Worth Metroplex area were enrolled. Women experiencing hot flashes took the survey at baseline and then 14 days, 2 months, and 6 months later. Factor analysis was performed. Face and content validity, internal consistency, test-retest reliability, and sensitivity to changes were evaluated. RESULTS: Fifty-two women (85.2%) completed all study sessions. The MVS survey was found to have good face and content validity and good internal consistency (rhoKR = 0.87). Spearman rank-order correlation coefficients at the 14-day retest varied from 0.58 to 1.00. As a result of these analyses, the MVS survey was further refined. CONCLUSIONS: The MVS survey was found to be comprehensive, simple, valid, reliable, sensitive, and a convenient instrument for subjective assessment of various characteristics of hot flashes. The MVS survey may serve as a valuable clinical and research tool for measurement of hot flashes.

Adult↗

A diagnostic tool for determining the quality of accuracy validation. Assessing the method for determination of nitrate in drinking water.

Realistic internal validation of a method implies the performance validation experiments under intermediate precision conditions. The validation results can be organized in an X (NrxNs) (replicates x runs) data matrix, analysis of which enables assessment of the accuracy of the method. By means of Monte Carlo simulation, uncertainty in the estimates of bias and precision can be assessed. A bivariate plot is presented for assessing whether the uncertainty intervals for the bias (E +/- U(E)) and intermediate precision (RSDi +/- U(RSDi) are included in prefixed limits (requirements for the method). As a case study, a method for determining the concentration of nitrate in drinking water at the official level set by 98/83/EC Directive is assessed by use of the proposed plot.

Bias↗

Maximizing internal and external validity in MMPI malingering research: a study of a military population.

The authors investigated the effectiveness of various commonly used Minnesota Multiphasic Personality Inventory (MMPI; Hathaway & McKinley, 1943) indices of exaggeration and malingering in detecting suspected malingering in a military sample of 121 enlisted men. To maximize external validity, only men undergoing psychological evaluation were used as participants. Forty-one participants were identified as suspected malingerers through multiple criteria and were contrasted with schizophrenic-spectrum and clinic outpatient groups. To improve internal validity, the 41 suspected malingering participants were asked to retake the test without exaggerating. Results revealed that there were many false positives and fewer, but nonetheless many, false negatives with standard malingering indices. It appeared that the Gough Dissimulation scale (Gough, 1947) might hold the most promise as a measure of malingering, but other scales are also useful. Individual comparisons between different samples and implications for MMPI-2 (Butcher et al., 1989) are presented.

Journal Article↗

Assessment of malingering with simulation designs: threats to external validity.

Comprehensive forensic evaluations are predicated on the accurate appraisal of response styles that may affect evaluatees' clinical presentation and experts' conclusions associated with psycholegal issues. In the assessment of malingering, forensic experts often rely heavily on standardized measures that have been validated exclusively via analogue research. While such research augments internal validity, the threats to external validity are readily apparent. As the first study of these threats, type of incentive (positive versus negative), context (a familiar versus unfamiliar scenario), and relevance to the participants was investigated systematically with a between-subjects factorial design. A sample of 231 undergraduates was asked to either (a) feign major depression and given an easily understood description of this disorder or (b) serve as controls responding honestly. They were administered a brief measure of psychopathology (Hopkins Symptom Checklist; Derogatis, Lipman, Rickels, Uhlenhuth, & Covi, 1974) and a recent screen for malingering (Screening Inventory of Malingered Symptoms or SIMS; Smith, 1992) in 1 of 18 experimental conditions. Results suggested that incentive had a main effect on the SIMS. More specifically, simulators under negative incentives appeared more focused in their feigning; they produced more bogus depressed symptoms, but fewer symptoms unrelated to depression. Interactions were also observed between context and incentive, and context and relevance. Implications of these results are explored for both analogue research on malingering and current forensic practice.

Adult↗

Criteria for evaluating the significance of developmental research in the twenty-first century: force and counterforce.

Since its birth approximately 100 years ago, the field of child development has undergone fluctuations in the criteria used to determine which research topics are more or less worthy of study. The purpose of this paper is to identify the forces that influence how developmental research is prioritized and evaluated and how these influences are changing as we enter the new millennium. We do so by considering the developmental researcher in context and suggest that there will be increasing pressure to use new criteria when assessing the significance of twenty-first-century developmental science. We review the three most commonly used forms of research validity--internal, external, and ecological--and then identify new research validities that we believe are likely to play increasingly important roles in the next millennium. We also argue that many developmental scientists will increasingly be pressured by forces that are external to the traditional research environment and that these forces will shape the ways in which the significance of developmental research is evaluated.

Child↗

Inference for the proportional hazards model with misclassified discrete-valued covariates.

We consider the Cox proportional hazards model with discrete-valued covariates subject to misclassification. We present a simple estimator of the regression parameter vector for this model. The estimator is based on a weighted least squares analysis of weighted-averaged transformed Kaplan-Meier curves for the different possible configurations of the observed covariate vector. Optimal weighting of the transformed Kaplan-Meier curves is described. The method is designed for the case in which the misclassification rates are known or are estimated from an external validation study. A hybrid estimator for situations with an internal validation study is also described. When there is no misclassification, the regression coefficient vector is small in magnitude, and the censoring distribution does not depend on the covariates, our estimator has the same asymptotic covariance matrix as the Cox partial likelihood estimator. We present results of a finite-sample simulation study under Weibull survival in the setting of a single binary covariate with known misclassification rates. In this simulation study, our estimator performed as well as or, in a few cases, better than the full Weibull maximum likelihood estimator. We illustrate the method on data from a study of the relationship between trans-unsaturated dietary fat consumption and cardiovascular disease incidence.

Biometry↗

Validity of indirect comparison for estimating efficacy of competing interventions: empirical evidence from published meta-analyses.

OBJECTIVE: To determine the validity of adjusted indirect comparisons by using data from published meta-analyses of randomised trials. DESIGN: Direct comparison of different interventions in randomised trials and adjusted indirect comparison in which two interventions were compared through their relative effect versus a common comparator. The discrepancy between the direct and adjusted indirect comparison was measured by the difference between the two estimates. DATA SOURCES: Database of abstracts of reviews of effectiveness (1994-8), the Cochrane database of systematic reviews, Medline, and references of retrieved articles. RESULTS: 44 published meta-analyses (from 28 systematic reviews) provided sufficient data. In most cases, results of adjusted indirect comparisons were not significantly different from those of direct comparisons. A significant discrepancy (P<0.05) was observed in three of the 44 comparisons between the direct and the adjusted indirect estimates. There was a moderate agreement between the statistical conclusions from the direct and adjusted indirect comparisons (kappa 0.51). The direction of discrepancy between the two estimates was inconsistent. CONCLUSIONS: Adjusted indirect comparisons usually but not always agree with the results of head to head randomised trials. When there is no or insufficient direct evidence from randomised trials, the adjusted indirect comparison may provide useful or supplementary information on the relative efficacy of competing interventions. The validity of the adjusted indirect comparisons depends on the internal validity and similarity of the included trials.

Data Interpretation, Statistical↗

Investigation of applicability of a mid-infrared spectroscopic method using an attenuated total reflection accessory and a new near-infrared transmission method for determination of faecal fat.

In many laboratories, the titrimetric method of Van de Kamer is used for the analysis of faecal fat content of patients suspected of steatorrhoea. We investigated the applicability of a mid-infrared (MIR) spectroscopic method, using an attenuated total reflection (ATR) accessory, and a new near-infrared (NIR) spectroscopic method. For the NIR method, sealed plastic bags containing the stool samples were used as transmission cells. Standardization was obtained using a previously described MIR method, with a NaCl flow-cell, as reference method. Partial least-squares regression was used for the calibration of each method. Full cross-validation of the calibration set was used for the internal validation of each method. Fifteen per cent of the stool samples could not be estimated with the ATR method within reasonable accuracy limits compared with the reference. The standard error of prediction of the NIR method was 1.1 g/dL. We conclude that the new NIR method is a promising technique for routine use. However, further experiments need to be done with triplicate measurements of each sample and the use of an external validation set.

Calibration↗