PubMed HealthSearch

SEARCH · PubMed Health

Results for “validity”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

[Validity of an ELISA test for CD4+ T lymphocyte count and validity of total lymphocyte count in the assessment of immunodeficiency status in HIV infection].

A newly available commercial ELISA (TRAx CD4, T Cell Diagnostics USA) for enumerating CD4+ T lymphocytes has been evaluated with blood samples of 105 HIV seropositive and 6 seronegative subjects. Results from the flow cytometric analysis were used as reference. The sensitivity and specificity of the ELISA to identify HIV seropositive subjects having less than 200 CD4+ T lymphocytes/microliters were assessed and studied using the ROC curve. The reproducibility of the ELISA test was analyzed on 40 samples. The results of the ELISA correlated well with these of the flow cytometric analysis (r = 0.79, p < 0.001). However, the ELISA test tends to overestimate the true CD4 count in HIV seropositives. This overestimation could not be explained by the aspecific contribution of monocytic CD4. The threshold for identifying HIV seropositive subjects with less than 200 CD4+ T lymphocytes with a maximum sensitivity and specificity was determined with ROC curve and equalled 400 cell equivalents with the ELISA (sensitivity and specificity were equal to 80%) and 1,450 lymphocytes/microliters with the total absolute lymphocyte count (sensitivity and specificity were equal to 75%). Using this curve, a threshold of 300 cell equivalents for the ELISA test and of 1,100 lymphocytes/microliters for the absolute lymphocyte count was shown to maximize the specificity (> 95%) without a significant loss of sensitivity.

CD4-Positive T-Lymphocytes

Measurement validity in physical therapy research.

This article considers the role of measurement validity within physical therapy research. The concept of measurement validity is identified as a component of internal validity, and it is differentiated from the notion of reliability; these concepts are related to systematic and random sources of error, respectively. Using examples from physical therapy and rehabilitation, four main types of validity are reviewed: face validity, criterion-related validity, content validity, and construct validity. The differing implications of these types of validity for quantitative and qualitative research are discussed. Three principal areas of concern are then addressed, based on a critical discussion of selected examples from the literature. First, it is argued that validity is often poorly distinguished from the allied concept of reliability and that purported claims for validity often only demonstrate reliability. Second, it is claimed that validity is too often neglected in favor of reliability, and specific examples relating to gait analysis are put forward to support this argument. Third, some of the methodological difficulties that may occur when attempts are made to demonstrate validity are considered. The article concludes with a plea for a closer focus on the issue of measurement validity within physical therapy research.

Bias

Externally validated risk prediction models for gestational diabetes mellitus: A systematic review and meta-analysis.

INTRODUCTION: Risk prediction models for gestational diabetes mellitus (GDM) offer potential for early identification and targeted prevention. External validation is crucial to assess model performance across diverse populations. Despite the availability of numerous GDM prediction models, limited evidence exists on their external validation frequency, methodological quality, and clinical applicability. This systematic review evaluated externally validated GDM prediction models, focusing on methodological rigor, reporting standards, and clinical relevance to inform future research and implementation. MATERIAL AND METHODS: Databases including Ovid MEDLINE, Embase, Scopus, Emcare, and CINAHL were searched up to May 1, 2025. Studies reporting external validation of GDM risk prediction models were included. Two reviewers independently screened studies. Data were extracted using the CHARMS framework, and risk of bias and applicability were assessed using PROBAST+AI. The study protocol was registered in the International Prospective Register of Systematic Reviews (PROSPERO; CRD420251125758). RESULTS: Twenty-six studies validated 33 models, with validation sample sizes ranging from 50 to 75&#x2009;161. Over half used the IADPSG criteria to define GDM. Discrimination metrics were commonly reported, but calibration, overall performance, and clinical utility were often lacking. Meta-analysis was feasible for only four models: Teede et&#xa0;al., Nanda et&#xa0;al., Naylor et&#xa0;al., and Van Leeuwen et&#xa0;al., each showing fair discrimination. The Teede et&#xa0;al. model was the most widely validated, with 11 external validations across six continents and a pooled AUC of 0.72 (95% CI: 0.67-0.76). Despite fewer validations, the Nanda et&#xa0;al. model achieved the highest pooled discrimination (5 validations; pooled AUC 0.77, 95% CI: 0.74-0.80). The Naylor et&#xa0;al. and van Leeuwen et&#xa0;al. models also underwent meta-analysis, as sufficient external validation studies were available to support comparative performance assessment. Notably, 69.23% of studies had a high risk of bias. CONCLUSIONS: While many models showed acceptable predictive performance, most validations were methodologically weak. Future studies should follow best-practice guidelines and promote scalable validation strategies, such as algorithm sharing, to enhance clinical utility.

Humans

Rational experimental design for bioanalytical methods validation. Illustration using an assay method for total captopril in plasma.

Generally, bioanalytical chromographic methods are validated according to a predefined programme and distinguish a pre-validation phase, a main validation phase and a follow-up validation phase. In this paper, a rational, total performance evaluation programme for chromatographic methods is presented. The design was developed in particular for the pre-validation and main validation phases. The entire experimental design can be performed within six analytical runs. The first run (pre-validation phase) is used to assess the validity of the expected concentration-response relationship (lack of fit, goodness of fit), to assess specificity of the method and to assess the stability of processed samples in the autosampler for 30 h (benchtop stability). The latter experiment is performed to justify overnight analyses. Following approval of the method after the pre-validation phase, the next five runs (main validation phase) are performed to evaluate method precision and accuracy, recovery, freezing and thawing stability and over-curve control/dilution. The design is nested, i.e., many experimental results are used for the evaluation of several performance characteristics. Analysis of variance (ANOVA) is used for the evaluation of lack of fit and goodness of fit, precision and accuracy, freezing and thawing stability and over-curve control/dilution. Regression analysis is used to evaluate benchtop stability. For over-curve control/dilution, additional to ANOVA, also a paired comparison is applied. As a consequence, the recommended design combines the performance of as few independent validation experiments as possible with modern statistical methods, resulting in optimum use of information. A demonstration of the entire validation programme is given for an HPLC method for the determination of total captopril in human plasma.

Calibration

Clinical Variable-Based Machine Learning for Predicting Early mCRPC Using Exclusively Clinical Variables: Development and Multicenter External Validation.

BACKGROUND AND OBJECTIVE: Metastatic hormone-sensitive prostate cancer (mHSPC) exhibits heterogeneous progression patterns, with early progression to metastatic castration-resistant prostate cancer (mCRPC) within 12 months indicating aggressive tumor biology and poor prognosis. Current risk stratification tools (CHAARTED, LATITUDE) offer limited individualized prediction. Machine learning approaches are increasingly applied to predict prostate cancer progression, but most models show modest performance (AUC 0.68-0.72), limited external validation, or require genomic variables unavailable in routine practice. This study aimed to develop and externally validate a novel RINH algorithm for predicting early mCRPC progression (&#x2264;&#x2009;12 months) using exclusively clinical variables, positioning it as a superior alternative to conventional ML classifiers. METHODS: This multicenter study enrolled 412 patients with de novo mHSPC from seven Spanish academic centers using mixed retrospective-prospective data collection. Twenty clinical variables were recorded, including demographics, PSA, ISUP grade, metastatic localization, CHAARTED/LATITUDE classifications, and treatment modalities. Following RINH-based outlier exclusion (55 patients), 357 patients (29 with early progression, 8.1%) were used to train six ML algorithms: RINH, Logistic Regression, Linear Discriminant, Support Vector Machine, Random Forest, and Subspace Discriminant. A two-tiered validation strategy integrated stratified fivefold cross-validation across all centers and formal external validation using center 1 (n&#x2009;=&#x2009;121, 19 events) for training and centers 2-7 (n&#x2009;=&#x2009;207, 10 events) for independent testing. Performance metrics included AUC, sensitivity, specificity, accuracy, and F1-score. KEY FINDINGS AND LIMITATIONS: Artificial intelligence and machine learning (ML) are transforming oncology, promising personalized risk stratification beyond traditional clinical criteria. In metastatic hormone-sensitive prostate cancer (mHSPC), early progression to castration resistance (mCRPC) within 12 months signals aggressive biology and poor prognosis, yet current tools (CHAARTED, LATITUDE) offer limited individualized prediction. Multiple ML models have been proposed with variable success: most achieve modest performance (AUC 0.68-0.72), lack robust external validation, or rely on genomic variables inaccessible in routine practice. We propose a novel approach using the Rivality Index Neighborhood (RINH) algorithm, demonstrating superior predictive capacity in an initial multicenter validation with exclusively clinical variables. This study provides rigorous multicenter external validation, advancing toward implementable precision oncology tools. CONCLUSIONS AND CLINICAL IMPLICATIONS: The RINH algorithm achieves superior predictive performance for early mCRPC progression using exclusively clinical variables, representing a significant advance toward implementable risk stratification. However, low reliability scores in external validation underscore that excellent performance metrics alone do not guarantee stability. Before clinical deployment, validation in substantially larger cohorts with higher progression events is essential. If validated, this model could enable personalized, risk-adapted therapeutic strategies, refining patient selection for treatment intensification or de-escalation.

Humans

[Reliability and validity of the Japanese version of the coping inventory for stressful situations (CISS): a contribution to the cross-cultural studies of coping].

OBJECTIVE: There has recently been a dramatic increase in the number of studies on coping behavior as an intervening variable between stress and health. Most of the available measures of coping are, however, psychometrically inadequate. We therefore decided to develop the Japanese version of the Coping Inventory for Stressful Situations (CISS) with special regards to its cross-cultural equivalence, reliability and validity. The CISS is a self-report measure of an individual's typical pattern of coping along three orthogonal dimensions of Task-, Emotion-, and Avoidance-oriented coping; its reliability and validity have been well studied in North America, where it was originally developed. METHOD: We obtained the Japanese version of the CISS (J-CISS) by means of back-translation. In Study 1, we administered the J-CISS and the 12-item General Health Questionnaire (GHQ) to 33 Japanese university students twice with an interval of four weeks. In Study 2,550 Japanese high school students completed the J-CISS and the Maudsley Personality Inventory. RESULTS: The equivalence of the Japanese version with the original was ascertained by means of back-translation involving multiple, independent mental health professionals and by factor congruence between the two versions. A principal component factor analysis (Varimax rotation) of the Study 2 data allowed us to extract three factors, which were virtually identical to the original ones. The high corrected item-remainder correlations, internal consistency reliabilities and test-retest reliabilities all attested to the reliability of the J-CISS. In order to examine its content validity, we compared the J-CISS with two coping questionnaires that have been in use in Japan, and found that the J-CISS covered most of the coping styles in these two questionnaires. However, such coping styles as "giving up," "to lose is to win (a Japanese proverb)," "it is best to do nothing" were not included in the original CISS and hence in the J-CISS. The criterion validity of the J-CISS was examined both in terms of predictive validity and concurrent validity. In Study 1, those who scored below the cut-off of the GHQ at Time 1 but above the cut-off at Time 2 had significantly higher Emotion-oriented coping scores at Time 1 than those who remained below the cut-off of the GHQ at Times 1 and 2 (predictive validity). In Study 2, the J-CISS scales and the MPI scales showed theoretically predicted correlations (concurrent validity). The results of the factor analysis and the corrected item-remainder correlations were suggestive of high construct validity of the J-CISS. Moreover, the mean inter-item correlation was between .20 and .40 for each scale, indicating its homogeneity. Factor analysis of each scale revealed that each scale indeed contained only one factor. Correlations among the three scales of the J-CISS established that the three scales formed multi-dimensional measures of coping. CONCLUSION: The results of our study indicate 1) that the obtained Japanese version of the CISS is to be regarded as final, 2) that coping styles can be measured in a consistent and reliable manner both in Japan and North America, and 3) that this cross-cultural equivalence as well as the other validity studies have further augmented the validity of the CISS itself.

Adaptation, Psychological

Evolution of a companywide validation program.

A successful computer validation program requires a solid foundation of policies, guidelines, and procedures. An unstructured approach to validation may yield adequate results for a specific computerized system, but it will not sustain ongoing, consistent computer validation activities over time. The evolution of a strong computer validation program starts with an awareness of the need for such a program and builds on that established base. One approach is a step-by-step method in which each new step builds on the results of one or more of the previous steps. However, the order of events is not as important as the ultimate completion of all the building blocks. A strong computer validation program requires a companywide policy and an administrator who provides oversight and chairs the computer validation committee. The committee should be established to provide departmental leadership and to develop company guidelines for computer validation. Committee members will also create system inventories and classify and prioritize those systems in terms of the need for validation. Departmental standard operating procedures must be developed to provide standardized methods for routine validation activities. Finally the quality assurance unit should have procedures in place that provide for its involvement in the computer validation process.

Computer Systems

Constructing and Validating Motive Bridging Inferences

Understanding Jane left early for the birthday party, She spent an hour shopping at the mall requires detecting that the first statement motivates the second. The validation model states that before accepting this bridging inference, the reader validates it with reference to relevant knowledge. In particular, a mediating idea is first derived from the text outcome and its candidate motive. If the mediating idea is supported by general knowledge, then the inference has been validated. In tests of this anaylsis, experimental subjects read motive or control sequences and then answered questions probing the knowledge hypothesized to validate the motive inferences, such as Do birthday parties involve presents? Five experiments confirmed that understanding motive sequences facilitates validating knowledge. A control procedure also refuted a priming counterexplanation of these effects (Experiment 1). Validation processing obtained for motive-outcome statements separated by two to four sentences in coherent sequences (Experiments 2 to 4). Inferred and explicit validating knowledge had a similar representational status (Experiment 3). Whereas proofreading abolished the validation effect, a reading strategy promoting causal processing did not enhance it (Experiment 4). A delayed priming procedure indicated that validating knowledge is integrated with the text representation (Experiment 5). The implications of these findings for the constructionist and minimal inference analyses were explored. The validation effects were simulated using construction-integration model.

Journal Article

Approaches for assessing the validity of a functional observational battery.

As neurobehavioral assessments during the preliminary stages of chemical testing are more widely undertaken, it is critical that the screening procedures utilized be valid indicators of neurobehavioral function and that they be sensitive, specific, and reliable. Efforts in this laboratory have been directed towards assessing these features in the use of a functional observational battery (FOB). For the purpose of assessing validity, we have examined FOB data which addresses the issues of criterion, predictive, concurrent, and construct validities. The FOB appears to be valid for detecting chemical-induced neurological dysfunction in rats, i.e., shows a good degree of criterion validity. Furthermore, in many instances the effects observed with the FOB may be predictive of symptomatology in humans. When comparisons can be made between effects detected with the FOB and other methods of measuring neurotoxicity (e.g., neuropathology), concurrent validity can also be established. To assess construct validity, effects of neurotoxicants can be classified into functional domains which are described by various measures in the FOB. Approaches for assessing the validity of the test method thus include answering specific research questions directed at assessing criterion, predictive, concurrent, and construct validity. Available data indicate that, in these aspects, the FOB is a valid screening method for the detection of neurotoxicity.

Animals

The validation of three human reliability quantification techniques--THERP, HEART and JHEDI: Part III--Practical aspects of the usage of the techniques.

This is the third paper in a series of three dealing with the detailed investigation of the empirical validity of three human reliability assessment (HRA) techniques. The first paper introduced the need for validation and specified the three techniques most requiring validation. The second paper detailed the results of an extensive independent validation experiment. This experimental validation involved 30 UK assessors using the techniques THERP, HEART and JHEDI (10 assessors per technique) to estimate the human error probabilities (HEPs) for 30 nuclear power and reprocessing (NP&R) tasks. The results for all three techniques were positive in terms of significant correlations, and general precision levels of 72% of all HEP estimates within a factor of 10 of the true value (unknown to the assessors). These results lend support to the empirical validity of these techniques in particular, and to HRA in general. However, the results were not all positive. In particular the consistency of usage of the techniques was variable. Additionally, subjects were generally not good at knowing their own uncertainty, i.e. they were not able to accurately predict when they were accurate nor when they were inaccurate. This desirable parameter is known as calibration, and the results from the validation suggested that subjects were not well-calibrated. This paper aims to determine how consistency of usage can be improved and to discern whether certain task types are, in practice, not well-assessed by the techniques, and hence are effectively currently beyond these techniques' abilities. Such information is aimed at aiding the HRA practitioner, or the ergonomist, interested in using these techniques. Recommendations for improving calibration are also discussed in this paper. A subsidiary but important focus of this paper is of a more fundamental nature, and of more general interest to the ergonomist. It concerns the validity of the techniques from an error reduction perspective. Currently these techniques may be used to identify how to reduce error probability, which is generally (in the qualitative sense) within the domain of ergonomics. One major mechanism for HRA-based error reduction is the utilisation of Performance Shaping Factor (PSF) information. This paper considers the validity of these PSF as ergonomics constructs. Drawing results from the validation exercise, it is seen how different PSF can be applied to the same scenario and can result in the same error probability, but will result in different error reduction guidance. It is therefore recommended that error reduction guidance must be based on a composite analysis of the results of the task, error identification and quantification analyses, with most weighting given to the qualitative analyses.

Evaluation Studies as Topic

Experience with two validation methods in a prevalence survey on nosocomial infections.

OBJECTIVE: To determine whether an investigator effect remained on the first German study on the prevalence of nosocomial infections Nosokomiale Infektionen in Deutschland Erfassung und Prävention (NIDEP), despite extensive validation efforts. DESIGN: Two validation methods were applied: bedside validation and validation by case studies. In both cases, the results of the four investigators were compared with the diagnosis of gold standard observers. SETTING: Validation measures were applied before, intermittently, during, and at the end of the surveillance period in 72 acute-care hospitals with 14,966 patients. RESULTS: The overall sensitivity in the bedside-validation periods was 89.0%; the overall specificity was 99.5%. For validation by case studies, overall sensitivity was 95.6%, and overall specificity was 92.8%. At the end of the surveillance, a remarkable investigator effect was found. CONCLUSION: Despite validation results that were assessed as satisfactory, based on available literature, an investigator effect was observed. This underlines the need for data validation and the formulation of recommendations for data validation. Clarification of the Centers for Disease Control and Prevention criteria for pneumonia and primary bloodstream infection and the inclusion of some diagnostic test results may reduce or prevent an investigator effect in future studies.

Bias