PubMed HealthSearch

SEARCH · PubMed Health

Results for “validation study”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The bogus pipeline as lie detector: two validity studies.

The validity of the bogus pipeline procedure for eliciting truthful responses from subjects in social psychological experiments was tested in two studies. Subjects were illegitimately informed about how to perform well on an experimental test, were tested, and then were asked whether they had possessed this prior information. As compared to subjects responding to pencil-and-paper questions or face-to-face questionning by the experimenter, those in a bogus pipeline condition confessed more often.

Electromyography

Artificial intelligence-based tumour infiltrating lymphocyte quantification in patients with triple-negative breast cancer: an independent validation study.

BACKGROUND: Tumour-infiltrating lymphocytes (TILs) are a robust prognostic marker in patients with triple-negative breast cancer. Artificial intelligence (AI)-derived computational tools assessing TILs could improve efficiency, but require independent validation against clinical outcomes. We aimed to compare the prognostic performance of AI-derived TIL scores with pathologist-scored TILs in a large, prospectively collected dataset pooled from randomised controlled trials. METHODS: CATALINA was an independent, external validation study using prospectively collected long-term clinical outcome data pooled from seven randomised clinical trials conducted at multiple sites. We independently evaluated two previously validated AI pipelines that generate five computationally assessed tumour-infiltrating lymphocyte (cTIL) scores by masked, independent deployment of locked models. cTIL scores were correlated with the mean of the pathologist-scored stromal TILs (sTILs) in 220 digitised haematoxylin and eosin whole slide images in a cohort of patients with early-stage triple-negative or HER-2 positive breast cancer, previously scored by trained pathologists in a TIL-reproducibility study. Prognostic performance was assessed in a separate cohort of patients with early triple-negative breast cancer pooled from seven prospective, randomised adjuvant trials. Multivariable Cox regression models adjusted for clinicopathological factors and study heterogeneity assessed associations of cTIL score and sTIL score with invasive disease-free survival, distant disease-free survival, and overall survival. 5-year discrimination was estimated using time-dependent area under the receiver operating characteristic curve (AUC). FINDINGS: Individual data were collated from 1759 patients, of whom 1356 had complete clinicopathological data, pathologist sTIL scores, and cTIL scores available. Modest correlation (r 0&#xb7;375-0&#xb7;473) was observed between cTIL scores and the mean pathologist sTIL score. Both sTIL and cTIL were independently associated with 5-year invasive disease-free survival, distant disease-free survival, and overall survival after adjustment for clinicopathological factors (hazard ratio for invasive disease-free survival was 0&#xb7;73 [95% CI 0&#xb7;66-0&#xb7;82]; q<0&#xb7;0001, distant disease-free survival was 0&#xb7;70 [0&#xb7;61-0&#xb7;79]; q<0&#xb7;0001, and overall survival was 0&#xb7;72 [0&#xb7;63-0&#xb7;82]; q<0&#xb7;0001 for sTIL scores and 0&#xb7;80 [0&#xb7;73-0&#xb7;89]; q<0&#xb7;0001, 0&#xb7;77 [0&#xb7;69-0&#xb7;86]; q<0&#xb7;0001, and 0&#xb7;79 [0&#xb7;70-0&#xb7;88]; q=0&#xb7;0002, respectively, for percentage_lymphocyte scores). In models adjusted for clinicopathological variables and sTIL score, cTIL score did not maintain a statistically significant prognostic association. Both sTIL and cTIL scores improved the 5-year AUC over clinicopathological variables alone, while cTIL score did not significantly further improve AUC when combined with clinicopathological variables and sTIL score. INTERPRETATION: Two cTIL models deployed entirely without retraining or modification provided statistically significant prognostic information and improved risk discrimination compared with clinicopathological variables alone in this large, platform-based, independent validation study. Although cTIL score did not incrementally improve prognostication compared with models combining clinicopathological variables with sTIL score, these findings support the application of cTILs as a reproducible prognostic biomarker, particularly in settings where routine or widespread pathologist assessment is unavailable. FUNDING: Breast Cancer Research Foundation (USA).

Humans

Measuring medical students' empathy: a validation study.

The aim of this study was to validate the empathy scale (Hogan, 1969) for use in the context of medical education in Australia. Empathy Scale scores of students in their first clinical year at Monash University were correlated with patient ratings, self ratings, and peer ratings of empathy. Inter-rater and intra-rater reliability were assessed. Correlations were also obtained between Empathy Scale scores and course marks in psychiatry. Of the empathy ratings only those by peers correlated significantly with Empathy Scale scores (r = 0-45, P less than 0-05, n = 29). Empathy Scale scores were unrelated to academic performance. In a separate part of the study, not connected to the establishing of criterion-related validity, Empathy Scale scores of the medical student group were found to be significantly higher (t=4-44, df = 52, P less than 0-001) than the scores of psychiatric patients with a diagnosis of "personality disorder". This study provides some support for the Empathy Scale as a measure of interpersonal effectiveness, but has not established it as a valid measure of empathy in a clinical setting.

Adult

Methods for defining equity-stratifying variables: a systematic review of validation studies.

BACKGROUND AND OBJECTIVE: Disease burden is often disproportionally higher among those who are socially disadvantaged by factors defined in the PROGRESS-Plus framework (ie, Place of residence, Race/ethnicity/culture/language, Occupation, Gender/sex, Religion, Education, Socioeconomic status, and Social capital, with "Plus" covering features like age and disability). The accuracy and applicability of case definitions to identify these variables from administrative and clinical health data are unknown. We conducted a systematic review to explore how equity-stratifying variables, as categorized by the PROGRESS-Plus framework, have been defined and validated in epidemiologic studies using administrative health, population-level, or electronic health record (EHR) data. METHODS: Medline, EMBASE, CINAHL, Web of Science, and Google Scholar were searched from the inception of the databases to 2024 for validation studies of equity-stratifying variables in adults using administrative health datasets, health registries, or EHR data. Titles and abstracts, followed by relevant full-text articles, were screened in duplicate by two reviewers for eligibility. The data sources utilized, algorithms employed, and their associated performance measures were extracted and synthesized from included studies. Given substantial heterogeneity in study design, equity-stratifying variable definition, and performance metrics, meta-analysis was not possible. RESULTS: Of the 9099 unique citations screened, 188 full texts were reviewed and 116 were included in this review. Most studies were published between 2019 and 2024 (n = 64, 55%) and were validation studies of race/ethnicity definitions that used race/ethnicity codes or surname list algorithms (n = 66, 57%). No studies examined religion. Regarding the reported performance measure estimates, the race/ethnicity/culture/language equity-stratifying variables category had the largest variability across sensitivity, positive predictive value (PPV), and Cohen's Kappa. Occupation validation studies had the lowest variation in sensitivity and PPV. CONCLUSION: Despite an increasing number of publications reporting on the validation of equity-stratifying variables relevant to the PROGRESS-Plus framework, performance measures varied widely across studies. The significant heterogeneity in equity-stratifying variable definitions and methods used to validate them support the need for further rigorous validation of equity-stratifying variables in administrative and clinical health data. PLAIN LANGUAGE SUMMARY: Disease burden is often higher in people who experience financial hardships, lower level of education, discrimination due to race/ethnicity, and unstable housing. These social factors can be considered health equity factors and are important for understanding health inequalities. Health researchers often use large datasets, such as hospital or electronic health records (EHRs), to study these health equity factors. However, it is not clear how accurately these data sources capture information about people's social circumstances and how these factors are defined. In this study, we reviewed existing research to understand how health equity factors have been defined across health data sources and how accurate they are at measuring aspects of health equity and social disadvantage. Of the more than 9000 studies we identified, we included 116 that met our criteria for this systematic review. Most included studies focused on identifying race and ethnicity, often using codes or surname-based methods. We found that the accuracy of these methods varied widely across studies, meaning results may not always be reliable or comparable. Overall, our findings show that there are inconsistencies in how social factors are defined and measured in health data. This makes it difficult to fully understand and address health inequalities using routinely collected health data. More work is needed to develop and validate better quality and more consistent methods for capturing these important social factors.

Humans

The Elizur test of psycho-organicity (adults). A cross-validation study.

In an effort to cross-validate the Elizur Test of Psycho-Organicity (performed on adults), a combined test used for the diagnosis of organic brain condition, 96 subjects were tested. The test can differentiate the "organic" and "non-organic" groups in a statistically significant manner. Eighty-five percent of the organics and 81% of the non-organics can be correctly identified by the test. In addition there is a significant correlation between the test findings and the electroencephalographic and radiographic results.

Adult

A validity study of the Neural Efficiency Analyzer in relation to selected measures of intelligence.

This study attempted to determine the validity of the Ertl Neural Efficiency Analyzer as a measure of intellectual ability by using NEA-Alpha and Neural Efficiency scores to predict collage grade point average (GPA) both alone and in combination with paper-and-pencil measures of intelligence for 22 male and 64 female college students. Results indicate that NEA-Alpha scores can predict GPA with moderate sucess and also that NEA-Alpha scores account for variability in grade point average not associated with paper-and-pencil tests.

Adult

Artificial intelligence-assisted histopathological diagnosis of endocervical gastric-type adenocarcinoma: a multicenter model development and validation study.

Endocervical gastric-type adenocarcinoma (GAS) is one of the most aggressive subtypes of cervical cancer and is frequently underdiagnosed due to morphological ambiguity, leading to delayed diagnosis. Despite the availability of molecular and genomic assays, their high cost, complexity, and limited reproducibility restrict clinical use. This study therefore proposes a highly sensitive artificial intelligence (AI)-assisted diagnostic system for GAS based exclusively on H&E-stained histopathological images. We included 309 slides from 96 GAS cases collected at Peking University Third Hospital from January 2018 to January 2025, representing the largest GAS cohort reported to date for AI research. In addition, we incorporated other morphologically analogous diseases, encompassing a total of 1,320 slides sourced from four categories: normal cervical mucosa (NORM), benign endocervical lesion entities (BELE), HPV-associated adenocarcinoma (HPVA), and endometrioid carcinoma with mucinous differentiation (ECMD). We developed GASPath, based on a novel multiple instance learning framework that efficiently captures fine-grained morphological variations from H&E-stained images. Beyond internal validation, GASPath was evaluated across 12 independent retrospective cohorts and further subjected to large-scale real-world validation on more than 7,000 samples from March 2024 to April 2025. Across three stages, GASPath demonstrated high performance. In internal validation (Stage I), it achieved an accuracy of 0.980 (95% CI 0.977-0.983) and an ROC-AUC of 0.995 (95% CI 0.994-0.997). In external validation (Stage II), the sensitivity reached 0.902 and improved to 0.968 with proposed strategies. For biopsy samples, GASPath achieved an ROC-AUC of 0.990 (95% CI 0.984-0.997). In large-scale real-world deployment (Stage III, n&#x2009;=&#x2009;7,056), GASPath achieved a balanced accuracy of 0.953, with 100% sensitivity for GAS (45/45 cases correctly identified). The heatmaps highlight morphological features of GAS that are easily underestimated, such as irregular, angulated glands, subtle loss of nuclear polarity, and mild cytologic atypia, which show substantial morphological overlap with other diagnostic categories. GASPath enables high-sensitivity detection of GAS in routine H&E-stained slides, obviating the need for extensive auxiliary testing while preventing underdiagnosis and misdiagnosis. This advancement addresses a critical gap by streamlining diagnostic workflows without compromising accuracy. Its implementation could enable cost-effective, scalable AI-assisted diagnostics, potentially transforming the early detection and management of this aggressive cancer subtype.

Female

Attachment behavior: a validation study in two age groups.

To assess the validity of attachment scores derived from the Ainsworth "strange situation," 56 1-year-olds and 79 2-year-olds accompanied by either the mother, the father, or a brief acquaintance were studied. Proximity to the adult, duration of play, crying, activity, and the incidence of looks and distance bids were measured. 1-year-olds were more secure with their parents: they were more active, played more, cried less, and stood closer to their parents than to an acquaintance. 2-year-olds accompanied by their parents were less settled in the presence of a stranger than children accompanied by the acquaintance. The adequacy of current conceptions and measures of attachment was discussed in light of these results.

Adult

Predicting 5-Year Mortality in Non-Small-Cell Lung Cancer Using the Korean Central Cancer Registry: Model Development and Validation Study.

BACKGROUND: Non-small-cell lung cancer (NSCLC) is one of the most common cancers and a leading cause of cancer-related mortality, making prognostic prediction clinically essential. Machine learning models are increasingly used to assess prognosis; however, developing systems that combine high discrimination with clear, clinically interpretable reasoning remains challenging. OBJECTIVE: This study aimed to develop deep learning models that predict 5-year mortality in NSCLC using data from the Korea Central Cancer Registry and quantify feature importance through permutation testing. METHODS: We identified 3144 patients diagnosed between 2014 and 2017 who had complete clinical data, pulmonary function test results, histological information, genomic data, and staging details. After preprocessing, the cohort was divided into stratified training, validation, and test sets in a 70%-15%-15% ratio. Five models were tuned using Hyperband across 10 predefined feature groups. The primary evaluation metric was the area under the receiver operating characteristic curve (AUC); additional metrics included accuracy, F1-score, precision, and recall. Groupwise permutation importance was calculated for each model, and the concordance of importance rankings was assessed using the Friedman test. RESULTS: All 5 models yielded comparable discrimination values on the test set (AUC=0.875-0.879). Model A was selected as the primary model and achieved an AUC of 0.879, an accuracy of 0.806, an F1-score of 0.824, and a Brier score of 0.142. Permuting the stage resulted in the largest decrease in AUC (0.217), followed by the pulmonary function test (0.016). Gene mutation had a modest overall impact but became more influential within the adenocarcinoma subset. The Friedman test showed no statistically significant differences in importance rankings across the models (P=.93). CONCLUSIONS: A grouped-input deep learning framework achieved discrimination comparable to a conventional Cox proportional hazards model using the same routine clinical variables for 5-year mortality prediction in NSCLC. Group-level permutation importance provided stable and reproducible insights into the clinical factors influencing risk, which may guide future model refinement and clinical decision-making.

Humans

Simultaneous determination of imiquimod and terbinafine in skin permeation studies: Validation of a liquid chromatography method with fluorescence detection.

Chromoblastomycosis is a chronic, neglected subcutaneous mycosis posing significant therapeutic challenges. A topical strategy combining terbinafine (TBF), an antifungal, with imiquimod (IMQ), a TLR-7/8 agonist immunomodulator, has emerged a promising alternative. However, no validated analytical method is currently available to simultaneously quantify both drugs in skin, which is crucial for novel formulation development. This study reports the development and validation of a simple HPLC method with fluorescence detection (excitation 236&#xa0;nm, emission 340&#xa0;nm) for the simultaneous determination of TBF and IMQ extracted from porcine skin. Separation was achieved on a C8 reversed-phase column (125&#xa0;&#xd7;&#xa0;4.0&#xa0;mm, 5&#xa0;&#x3bc;m) using a mobile phase of methanol and water (60,40, v/v), both containing 0.1% formic acid at a flow rate of 0.8&#xa0;mL/min. The method showed excellent linearity (r&#xa0;>&#xa0;0.999) over 0.01-1.0&#xa0;&#x3bc;g/mL for IMQ and 0.1-2.0&#xa0;&#x3bc;g/mL for TBF. Intra- and inter-day precision demonstrated coefficients of variation below 5%, and recovery rates from skin (79-105%) confirmed accuracy. Limits of detection were 0.001&#xa0;&#x3bc;g/mL for IMQ and 0.004&#xa0;&#x3bc;g/mL for TBF, with quantification limits of 0.02&#xa0;&#x3bc;g/mL and 0.16&#xa0;&#x3bc;g/mL, respectively. This selective, sensitive, and reproducible method represents a valuable analytical tool for supporting the development and quality control of topical formulations for chromoblastomycosis and other fungal skin diseases.

Animals

A cross-validation study for predictors of scores on state board examinations.

Regression equations were computed and cross-validated for predicting nurses' scores on each of the five state board examinations (SBEs). Scores on five National League for Nursing (NLN) achievement tests of 101 nurses who graduated from a baccalaureate program between 1968 and 1972 were used as predictor variables. Stepwise regression analysis was used to compute the equations which were, in turn, used to predict the SBE scores for an independent sample of nurses who graduated in 1973. Cross-validation correlations between the predicted and actual obtained scores of the 1973 graduates ranged from .64 to .81. For each equation, the NLN test scores in Nursing of Children and Obstetric Nursing were consistently the best indicators of performance on the SBEs. Additionally, a factor analysis indicated that the SBEs do not measure independent entities, but that all five examinations have high loadings on the same factor. The results suggest that nursing programs should develop and validate prediction equations to assist nurses in preparing for SBEs.

Achievement

Reliability and validity studies of the comprehensive Psychopathological Rating Scale (CPRS).

1. The most important properties of a rating instrument aimed at detecting possible changes in psychopathology are as follows: a) communicability to not yet experienced raters; b) satisfactory levels of reliability and validity; and c) sensitivity to changes. Since its construction some years ago these properties of the Comprehensive Psychopathological Rating Scale (CPRS) have been investigated in several sessions, both experimental, and in the course of clinical drug trials. 2. Raters with a different training background (physicians, psychologists, social workers, nurses), and a different level of experience in ratings have participated in the sessions. 3. In all instances the CPRS has proved to be highly reliable and highly sensitive. It appeared to be easily communicable, also in its international versions, and its content seems to fit very well with different cultural contexts. Information about its validity has been inferred from the studies mentioned above, and, also from correlations with other well known rating instruments.

Humans

The Middlesex Hospital Questionnaire: a validity study.

The MHQ is a brief self-rating inventory purporting to measure aspects of six distinct categories of psychoneurosis and affective status. It has been found to be a reliable instrument and also valid as a profile measure. Two individual scales have also previously been explored in respect of validity. The present report describes a further attempt to examine the validity of individual scales in relation to pertinent single clinical diagnostic entities in a study involving 800 patients. The phobic and obsessional scales are found to be particularly accurate and differentiating in this respect. Patients variously diagnosed as suffering from anxiety states, depressive states and personality disorder tend to score very highly on several scales. The instrument serves overall to distinguish satisfactorily between such populations and others suffering from schizophrenia and anorexia nervosa. It also markedly differentiates them from 'normal' populations.

Anorexia Nervosa

Assessing depressive symptoms in five psychiatric populations: a validation study.

Data from five psychiatric populations and a community sample are presented on the CES-D, 20-item self-report depression symptom scale developed by the Center for Epidemiologic Studies. Results show that the scale is a sensitive tool for detecting depressive symptoms and change in symptoms over time in psychiatric populations, and that it agrees quite well with more lengthy self-report scales used in clinical studies and with clinician interview ratings. Although a symptom scale cannot differentiate between diagnositc groups, the CES-D has demonstrated its validity as a screening tool for detecting depressive symptoms in psychiatric populations.

Adolescent

A validational study of the WIST as a group-administered instrument for assessment of schizophrenic thinking.

The objective assessment of thought disorder in schizophrenia is problemmatic in clinical psychology. Recently an individually administered instrument (WIST) was introduced as a brief, objective, and quantitative measure of schizophrenic thought processes. Possible shortcomings of the WIST are noted; experimental findings that concern extension to group testing conditions, convergent validity with another self-report measure of schizophrenia, and discriminant validity from intellectual level are presented.

Adult

An individualized nomogram for predicting progression-free survival in systemic anaplastic large cell lymphoma: a multicenter, retrospective, and internally validated study.

OBJECTIVES: To develop an individualized nomogram for predicting disease progression risk in systemic anaplastic large cell lymphoma (sALCL). METHODS: Independent predictors of progression-free survival (PFS) were identified using Cox regression in a multicenter retrospective cohort of 109 sALCL patients (2010-2022). These were incorporated into a three-factor nomogram, evaluated via bootstrapped internal validation (1000 resamples), ROC analysis, C-index, decision curve analysis (DCA), and clinical impact curve (CIC). RESULTS: A total of 29 PFS events occurred during a median follow-up of 31 months. Multivariable modelling selected serum &#x3b2;2-microglobulin elevation, extranodal disease, and front-line chemotherapy choice (CHOP versus CHOPE or BV+CHP) as autonomous progression drivers. Upon internal bootstrap validation, the nomogram yielded strong prognostic accuracy, achieving AUCs of 0.81, 0.85 and 0.87 for 1-, 3- and 5-year progression-free survival, alongside a corrected C-index of 0.779 (95% CI: 0.699 - 0.861). Calibration plots showed close agreement between predicted and observed outcomes, while DCA confirmed superior net clinical benefit versus conventional IPI or Ann Arbor stratification across multiple decision thresholds. CONCLUSION: This first sALCL-specific nomogram integrates clinical and treatment variables to provide personalized PFS risk estimation. While internally validated, this exploratory, observation-based tool requires external validation and recalibration in prospective cohorts before clinical implementation.

Humans

The measurement of obsessionality: first validation studies of the Lynfield obsessional/compulsive questionnaires.

The difficulties encountered in attempts at the rating and quantification of obsessional behaviour are outlined and some of the reasons for such difficulties are mentioned. The development of the Lynfield Obsessional-Compulsive Questionnaire from the Leyton Obsessional Inventory is discussed together with the advantages and disadvantages attendant upon its administration. The author concludes that the Lynfield Obsessional-Compulsive Questionnaire promises well as a measure for assessing change and that, following further validity and reliability tests, it may be used in multicentre trials of treatment procedures.

Evaluation Studies as Topic

Dream reports and the test of emotional styles: a convergent - discriminant validity study.

Thirty-eight undergraduates completed the forced-choice form of the Test of Emotional Styles and participated in a 4-day dream recall study. Correlational analysis was performed with the three dimensions of emotional style and two dream report aspects as variables. The dimensions failed to correlate significantly with any of the dream report variables, although the dimensions intercorrelated significantly among themselves. Doubt is expressed as to the construct validity of the subscales of the Test of Emotional Styles.

Affect