PubMed HealthSearch

SEARCH · PubMed Health

Results for “validity”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

[Validity of the Hoppe List. An empirical study of patients with backache].

Construct validity, criterion-oriented validity, and differential validity were assessed to determine the validity of the Hoppe-Liste (HL). Construct validity was ascertained by the multitrait-multimethod analysis suggested by Campbell & Fiske. Data regarding both convergent and discriminant validity show that the HL measures the construct pain. A comparison of the HL with the Revised Multidimensional Pain Scale and the Visual Analogue Scale suggests the criterion-oriented validity of the HL. Group-specific scores of the HL scales and of correlation patterns between scales tentatively support differential validity. It is concluded that the HL seems to be a suitable instrument for measuring the intensity and quality of pain.

Back Pain

Validation of a Turkish Translation of the Stress in Emergency Healthcare Professionals: The Stress Factors and Manifestations Scale.

AIM: The primary duties of emergency healthcare professionals (EHPs) are to provide emergency patient care to acutely ill and injured individuals. Due to the nature of their work, EHPs operate under constant stress, often requiring rapid decision-making, swift action, and the delivery of necessary medical care in life-or-death situations, sometimes under inadequately safe conditions. Therefore, the aim of this study is to determine the validity and reliability of the Emergency Healthcare Professional Stress Factors and Symptoms (SEHP:SFMS) Scale in Turkish for identifying stress factors and symptoms in emergency medical care professionals providing emergency patient care services. DESIGN: A methodological study design was used in this study. METHODS: The study was conducted with the participation of 211 EHPs from employees working in emergency care institutions affiliated with the Muğla Provincial Health Directorate between November 2023 and June 2024. Data were collected via a face-to-face survey. Data were analysed using Lawshe content validity ratio, Kaiser-Meyer-Olkin coefficient, Bartlett test, exploratory factor analysis, principal component analysis, Varimax factor rotation method, confirmatory factor analysis, Cronbach's α internal consistency coefficient, convergent validity, discriminant validity, test-retest, and Spearman correlation coefficient tests. RESULTS: The linguistic translation and cultural adaptation of the SEHP:SFMS showed strong performance. The scope validity index of the scale is 0.83. The item-total correlation values of the scale were found to be between 0.486 and 0.794, and the factor loadings were between 0.474 and 0.816. Confirmatory factor analysis fit indices: χ2 = 248.727; df = 101; n = 211; p = 0.000; χ2/df = 2.463; RMSEA = 0.083; CFI = 0.914, SRMR = 0.052, which was found to be compatible and acceptable with the proposed 3-factor model. The Cronbach's α reliability coefficient of the scale was 0.931, and the total variance was 61.97%. CONCLUSIONS: SEHP:SFMS is a valid and reliable tool to assess stress factors and symptoms of Turkish emergency healthcare professionals. Its use improves the quality of emergency care. PATIENT OR PUBLIC CONTRIBUTION: These study findings have been used to create a tool with Turkish validity and reliability that allows for the examination of stress factors among healthcare professionals working in emergency and critical services. Identifying and reducing stress factors among healthcare professionals is crucial for the delivery of quality healthcare services. It can also be used to develop targeted interventions and ongoing strategies to facilitate improved clinical supervision and mentoring. IMPLICATION FOR NURSING PRACTICE: Nurses in emergency departments, which are among the most stressful, dynamic, intense, life-saving, and critical environments in healthcare institutions, and where life-saving treatment is administered, are at high risk of experiencing psychological trauma. Trauma experienced in the work environment is a significant problem for nursing. The consequences of trauma negatively affect nurses and institutions. Studies show that post-traumatic stress, anxiety, depression, and burnout are commonly observed in emergency department nurses. In this sense, understanding the stress and stress factors experienced by nurses can guide future interventions. The results of this study are considered important in making visible the stress and stress factors experienced by nurses in the emergency department, and also in guiding managers and nurses working in this field in terms of preventive and protective measures.

Humans

Reliability and descriptive validity of PSE syndromes.

Despite extensive research use of the Present State Examination (PSE), the validity of classification based on PSE data has not been studied extensively. We have examined a consecutive series of functional psychiatric admissions using the PSE and systematically gathered clinical and demographic data in order to study not only the reliability but the descriptive (construct) validity of classification based on PSE data. We have found that the PSE can be used in a psychiatric hospital to reliably describe and classify schizophrenic and affective syndromes with considerable descriptive validity in terms of clinical and demographic variables. We believe that this type of validity is an important step in establishing validity of clinical disease entities. The interrelationship among different kinds of validity (descriptive, concurrent, predictive) might provide clinical disease concepts with more definitive validity.

Adult

Analytical methods validation: bioavailability, bioequivalence and pharmacokinetic studies. Conference report.

This is a summary report of the conference on Analytical Methods Validation: Bioavailability, Bioequivalence and Pharmacokinetic Studies. The conference was held from December 3 to 5, 1990 in the Washington, DC area and was sponsored by the American Association of Pharmaceutical Scientists, US Food and Drug Administration, Federation International Pharmaceutique, Health Protection Branch (Canada) and Association of Official Analytical Chemists. The purpose of the report is to represent our assessment of the major agreements and issues discussed at the conference. The report is also intended to provide guiding principles for validation of analytical methods employed in bioavailability, bioequivalence and pharmacokinetic studies in man and animals. The objectives of the conference were: 1. To reach a consensus on what should be required in analytical methods validation and the procedures to establish validation; 2. To determine processes of application of the validation procedures in the bioavailability, bioequivalence and pharmacokinetic studies; 3. To develop a report on analytical methods validation (which may be referred to in developing future formal guidelines). Acceptable standards for documenting and validating analytical methods with regard to processes, parameters or data treatments were discussed because of their importance in assessment of pharmacokinetic, bioavailability and bioequivalence studies. Other topics which were considered essential in the conduct of pharmacokinetic studies or in establishing bioequivalency criteria, including measurement of drug metabolites and stereoselective determinations, were also deliberated.

Biological Availability

Issues of validity in the Diagnostic Interview Schedule.

The Diagnostic Interview Schedule, the chief instrument in contemporary studies in psychiatric epidemiology, enhances the reliability of psychiatric diagnosis and enables lay interviewers to closely reproduce psychiatric interviews. However, despite frequent references in the literature to the validity of the Diagnostic Interview Schedule, most studies fundamentally represent variations of reliability paradigms to the neglect of criterion-related validity. Mistaken assertions of validity persist in the psychometric language used to describe the Diagnostic Interview Schedule. This article examines the basis for claims and counterclaims of validity in accordance with standard psychometric definition, and identifies sources of erroneous reasoning in attempts to infer validity from reliability. The article presents a general framework organizing the process of diagnostic validation and discusses strategies for research seeking to validate psychiatric diagnoses achieved through the Diagnostic Interview Schedule.

Humans

Automating candidate gene prioritization with large language models: from naive scoring to literature-grounded validation.

MOTIVATION: Identifying promising therapeutic targets from thousands of genes in transcriptomic studies remains a major bottleneck in biomedical research. While large language models (LLMs) show potential for gene prioritization, they suffer from hallucination and lack systematic validation against expert knowledge. RESULTS: The framework identified 609 sepsis-relevant genes with >94% filtering efficiency, demonstrating strong enrichment for inflammatory pathways including TNF-α signaling, complement activation, and interferon responses. Literature validation yielded 30 ultra-high confidence therapeutic candidates, including both established sepsis genes (IL10, TREM1, S100A9, NLRP3) and novel targets warranting investigation. Benchmark validation against expert-curated databases achieved 71.2% recall, with systematic correlation between computational confidence and evidence quality. The final candidate set balanced discovery (11 novel genes) with validation (19 known genes), maintaining biological coherence throughout the filtering process. This framework demonstrates that rigorous methodology can transform unreliable LLM outputs into systematically validated biological insights. By combining computational efficiency with literature grounding, the approach provides a practical tool for prioritizing experimental validation efforts. The modular design enables adaptation to other diseases through knowledge base substitution, offering a systematic approach to literature-guided biomarker discovery. AVAILABILITY AND IMPLEMENTATION: We developed a two-stage computational framework that combines LLM-based screening with literature validation for systematic gene prioritization. Starting with 10 824 genes from the BloodGen3 repertoire, we applied multi-criteria evaluation for sepsis relevance, followed by retrieval-augmented generation using 6346 curated sepsis publications. A novel faithfulness evaluation system verified that LLM predictions aligned with retrieved literature evidence. Source code and implementation details are available at https://github.com/taushifkhan/llm-geneprioritization-framework, vector database at https://doi.org/10.5281/zenodo.15802241, and Interactive demonstration at https://llm-geneprioritization.streamlit.app/.

Humans

Development and content validity testing of a comprehensive classification of diagnoses for pediatric nurse practitioners.

Pediatric nurse practitioners (PNPs) need an integrated, comprehensive classification that includes nursing, disease, and developmental diagnoses to effectively describe their practice. No such classification exists. Further, methodologic studies to help evaluate the content validity of any nursing taxonomy are unavailable. A conceptual framework was derived. Then 178 diagnoses from the North American Nursing Diagnosis Association (NANDA) 1986 list, selected diagnoses from the International Classification of Diseases, the Diagnostic and Statistical Manual, Third Revision, and others were selected. This framework identified and listed, with definitions, three domains of diagnoses: Developmental Problems, Diseases, and Daily Living Problems. The diagnoses were ranked using a 4-point scale (4 = highly related to 1 = not related) and were placed into the three domains. The rating scale was assigned by a panel of eight expert pediatric nurses. Diagnoses that were assigned to the Daily Living Problems domain were then sorted into the 11 Functional Health patterns described by Gordon (1987). Reliability was measured using proportions of agreement and Kappas. Content validity of the groups created was measured using indices of content validity and average congruency percentages. The experts used a new method to sort the diagnoses in a new way that decreased overlaps among the domains. The Developmental and Disease domains were judged reliable and valid. The Daily Living domain of nursing diagnoses showed marginally acceptable validity with acceptable reliability. Six Functional Health Patterns were judged reliable and valid, mixed results were determined for four categories, and the Coping/Stress Tolerance category was judged reliable but not valid using either test. There were considerable differences between the panel's, Gordon's (1987), and NANDA's clustering of NANDA diagnoses. This study defines the diagnostic practice of nurses from a holistic, patient-centered perspective. It is the first study to use quantitative methods to test a diagnostic classification system for nursing. The classification model could also be adapted for other nurse specialties.

Humans

Validity of repeated dietary measurements in a dietary intervention study.

The aim of the study was to evaluate the compliance in a dietary intervention study. When drawing conclusions about the relationship between dietary intake and disease occurrence/disease-related variables it is important to obtain valid dietary data. 20 healthy, non-smoking normal-weight omnivores changed from a mixed to a lactovegetarian diet. Dietary surveys (four 24 h recalls per person and time-period), urinary and faecal sample collections were performed before and 3, 6 and 12 months after the dietary shift. The validation of energy, protein, sodium and potassium yielded approximately the same ratio of dietary intake to biological marker at 0 and 3 months. This ratio decreased towards 6 months and continued to decrease towards 12 months. The fibre intake was compared to the total faecal weight directly and indirectly by calculating the fibre intake from the stool weight, the water content in faeces and the excretion of short-chain fatty acids (SCFAs). These four methods of fibre validation showed that the ratio of dietary intake to biological marker was always highest at 12 months, indicating an overestimation of the fibre intake at the end of the study. This is the first time these methods of validating fibre intake have been used in an epidemiological study. The ratio of dietary calcium intake to urinary and faecal calcium excretion did not show any statistical difference between the period before and 3 months after the dietary shift. To conclude, almost all investigated dietary data show approximately the same validity before and 3 months after the dietary shift, and show the least validity 12 months after the dietary shift. Thus, this study demonstrates that it is difficult to obtain valid dietary data 1 year after a drastic dietary change, indicating a decreased compliance to the new dietary regimen at the end of the 1 year study period. This represents important information when attempting to relate biological effects to dietary intake, and illustrates the importance of using biological markers for food intake in dietary surveys.

Adult

[Adaptation and validation of a test on knowledge about diabetes mellitus].

OBJECTIVE: To adapt and validate a Spanish language medium test, of theoretical general knowledge of diabetes mellitus (questionnaire from the University of Michigan). To determine the validity of the concurrent and discriminatory content and establish reliability. DESIGN: The study was observational. Validity was verified prior to data collection. To analyse the concurrent and discriminatory validity, a questionnaire was used in personal interview with the patients, and the degree of knowledge evaluated on certain variables. SETTING: Hospital outpatient endocrinology consultations. PATIENTS: 167 diabetic patients were chosen at random, from the outpatient visits. 14 patients who had developed hearing, or language problems, or who had problems of a psychological nature, were excluded. Only 1 patient refused to answer the questionnaire. MAIN MEASUREMENTS AND RESULTS: Validity of the content was confirmed after careful analysis of the questions on the questionnaire by medical specialists in endocrinology. It was found that the test had adequate concurrence (p less than 0.01) when the average general knowledge levels of certain group of patients are compared. It also had acceptable discriminatory validity (r = 0.56: p less than 0.0001) and reliability (alpha: 0.84; p less than 0.45). CONCLUSIONS: Adaptation and validation has been obtained for a test of theoretical general knowledge on diabetes mellitus, and the test was found to be applicable to the population under study.

Adult

[Validation in the French language of an evaluation scale for family functioning (FACES III): a tool for research and for clinical practice].

Valid and reliable scales are necessary to describe the numerous factors associated with health status. Multiple studies have shown the impact of familial factors, i.e. family functioning (FF) on health indicators. As no scale exists in french to assess FF, we performed a study to validate the Family Adaptability and Cohesion Evaluation Scale (FACES III, Olson et al.) in french. This scale has a high level of validity and reliability and has been widely used. After 2 translations and back-translations and a pilot study to establish face validity, the final version was studied in 976 healthy subjects (457 families) who attended a preventive medical center in Nancy, France. Parents and adolescents each filled out two 20 item self-administered questionnaires. There were few missing values (0-3% per item). Construct validity was assessed by principal component factor analysis, which found the same two individualized axes as in the original scale. The reliability of the french version was excellent and comparable to Olson's scale. This study demonstrates the validity and reliability of a french version of FACES III in a french population. It provides researchers and clinicians in France with a validated instrument for assessing, in a quantified way, factors associated with family functioning that influence health status in adults, adolescents and children, particularly those with chronic diseases. This scale is especially useful for developing and evaluating health programs.

Adolescent

Catecholaminergic polymorphic ventricular tachycardia mediated by ryanodine receptor 2: a validated risk stratification.

BACKGROUND AND AIMS: Patients with catecholaminergic polymorphic ventricular tachycardia (CPVT) are at risk for potentially life-threatening arrhythmic events (AEs) even while treated with β-blockers. The aim was to develop a model for individualized prediction of AEs in patients with RYR2-mediated CPVT on β-blocker monotherapy. METHODS: The derivation and independent validation cohorts included 743 and 129 patients, respectively. AEs were defined as arrhythmic syncope, appropriate implantable cardioverter-defibrillator shock, sudden cardiac arrest (SCA), and sudden cardiac death. Near-fatal or fatal AEs (nf/fAEs) included all AEs except for arrhythmic syncope. Prediction models using Cox regression were developed and internally and externally validated. RESULTS: A total of 102 (13.7%) patients in the derivation cohort and 24 (18.6%) patients in the validation cohort experienced ≥1 AE over a median follow-up of 5.1 [interquartile range (IQR), 7.7] and 2.4 (IQR, 4.4) years, respectively. Predictors of AE were arrhythmic syncope or SCA prior to diagnosis and age at β-blocker initiation. In the derivation and validation cohorts, the optimism-corrected C-indices of the models for AE were 0.67 [95% confidence interval (CI) 0.62-0.72] and 0.59 (95% CI 0.48-0.71), respectively. For nf/fAEs, ventricular arrhythmia severity before β-blocker initiation was a fourth independent predictor, and C-indices of the models in the derivation and validation cohorts were 0.74 (95% CI 0.68-0.80) and 0.60 (95% CI 0.47-0.72), respectively. In the derivation cohort, calibration slopes were 1.00 (95% CI 0.59-1.41) for AE and 1.00 (95% CI 0.69-1.32) for nf/fAE. CONCLUSIONS: These externally validated risk prediction models using clinical parameters accurately distinguished CPVT patients on β-blocker monotherapy at low and high risk for future AEs while treated with β-blockers. These models provide guidance for implementation of clinical management therapies to prevent AEs in patients with CPVT.

Humans

Assessing the Concurrent Validity of the Australian Treatment Outcomes Profile in a Methamphetamine Dependent Treatment-Seeking Population.

INTRODUCTION: The Australian Treatment Outcomes Profile (ATOP) is a brief clinical tool assessing substance use, health and well-being used in Australian alcohol and other drug treatment services. It is validated for use with clients using alcohol, opioids and cannabis, but not yet for clients who primarily use methamphetamine. METHODS: An embedded validation study was undertaken in treatment-seeking adults enrolled in a randomised double-blind placebo-controlled trial of lisdexamfetamine for methamphetamine dependence with sites in New South Wales, South Australia and Victoria. Participant demographics were collected during study screening. The ATOP and comparators (Time Line Follow Back, Opiate Treatment Index, Depression Anxiety Stress Scale, WHOQOL-BREF and Personal Wellbeing Index) were collected at baseline. Continuous ATOP items were analysed using Pearson's correlation coefficient, and dichotomous items were analysed using Fleiss's &#x3ba;. Agreement was rated as strong where measures were &#x2265;&#x2009;0.50, moderate where agreement was 0.30-0.49, and weak where <&#x2009;0.30. RESULTS: One hundred and eighteen study participants (2018-2020) had data for concurrent validity analysis. Strong validity was demonstrated for physical health, psychological health, quality of life, injecting drug use and crime items, and for days of use for amphetamines, alcohol, cannabis and cocaine. There was weak validity for days of use for benzodiazepines. Heroin use days and other opioid use days were endorsed by fewer than five participants and were therefore unable to be assessed. DISCUSSION AND CONCLUSIONS: The ATOP is valid for use in a treatment-seeking methamphetamine-dependent population, expanding the range of tools for assessment and standardised outcome monitoring across different settings and services.

Humans

Discovery and validation of GNA12circle as a first-trimester plasma eccDNA marker for early-onset preeclampsia.

BACKGROUND: Early-onset preeclampsia (EOPE) is a major cause of maternal and perinatal morbidity and is characterized by placental dysfunction, systemic endothelial injury, and hypertensive vascular stress. Because hypertensive disorders of pregnancy may also signal later maternal cardiovascular and cerebrovascular vulnerability, effective biomarkers for first-trimester risk assessment remain clinically important. Extrachromosomal circular DNA (eccDNA), a stable form of circulating cell-free DNA, has emerged as a potential source of disease-associated biomarkers. This study aimed to characterize first-trimester plasma eccDNA alterations associated with subsequent EOPE and to identify and validate a candidate circulating eccDNA marker for early risk assessment. METHODS: A two-stage nested case-control study was conducted within a prospective birth cohort. In the discovery stage, plasma samples collected at 11-13&#x202f;weeks of gestation from 5 women who subsequently developed EOPE and 5 matched normotensive controls were profiled by Circle-Seq to characterize genome-wide eccDNA alterations. Candidate eccDNAs were prioritized through differential abundance analysis and were further confirmed by outward PCR and Sanger sequencing. In the validation stage, the candidate selected marker was quantified by junction-specific qPCR in an independent cohort of 109 EOPE cases and 109 controls. Its potential predictive value was further evaluated alone and in combination with routine first-trimester clinical variables. RESULTS: In the exploratory discovery analysis, 410 nominally differentially abundant candidate eccDNAs were identified as a hypothesis-generating pool. Among these, GNA12circle (chr7:2876332-2,876,692) was prioritized and experimentally validated at the circular junction. In the independent validation cohort, plasma GNA12circle levels were significantly higher in women who later developed EOPE than in controls. When combined with routine first-trimester variables, GNA12circle improved predictive performance. The RF model showed the best overall cross-validated performance among the evaluated classifiers, with a mean held-out test-fold AUC of 0.843. CONCLUSION: First-trimester plasma eccDNA profiling revealed distinct alterations associated with subsequent EOPE, from which GNA12circle was identified and validated as a candidate circulating marker. These findings support further investigation of circulating eccDNA for early EOPE risk assessment in larger multicenter populations.

Humans

Trustworthy Agentic AI in Bioinformatics: From Workflow Automation to Traceable and Validated Biological Inference.

Agentic artificial intelligence is extending bioinformatics beyond conversational assistance by enabling systems to select tools, execute code, revise analytical plans, and interpret biological data. These capabilities may accelerate research, but they also redistribute decisions that determine whether biological conclusions are valid. We conducted a targeted, structured PubMed search in July 2026 and identified 11 peer-reviewed agentic bioinformatics systems for descriptive review based on predefined eligibility criteria for analytical decision-making, tool or code execution, iterative evaluation, or coordinated agent activity. The evidence base covered single-cell transcriptomics, microbial genomics, cancer genomics, and omics applications, together with methodological literature on reproducibility and biological validation. We examined how current systems report delegated authority, provenance, validation, evidence, abstention, and human oversight. Existing platforms implement safeguards such as sandboxed execution, restricted commands, interaction logs, evidence identifiers, automated checks, critic agents, quality scores, and expert assessment. However, published reports rarely provide a connected account linking the original biological question to samples, reference resources, analytical decisions, computational actions, statistical results, supporting evidence, validation outcomes, and final claims. We distinguish inherited bioinformatics errors, errors amplified through autonomous action, and emergent failures arising from memory, retrieval, tool interaction, or agent coordination. We further propose a multidimensional decision-rights profile, consequence-sensitive validation gates, and a claim-to-evidence provenance architecture organized through the Traceable History of Research Evidence, Agent Actions, and Decisions in Bioinformatics (THREAD-Bio) framework. Illustrative cases show that technically successful execution may still support misleading inference. Trustworthy agentic bioinformatics therefore requires claims to remain reconstructible, challengeable, validated, and proportionate to the evidence.

accountable autonomy

Reliability and validity in binary ratings: areas of common misunderstanding in diagnosis and symptom ratings.

Confusion may exist between the reliability of a binary rating (for example, schizophrenia versus not-schizophrenia) and its implications for validity. High reliability does not guarantee validity, but paradoxically, low reliability does not imply poor validity in all contexts. Changes in the base rate or in experimental design may indicate high validity even when the reliability was thought to be low. Attempts to improve the psychiatric nomenclature by increasing only reliability run the risk of the "attenuation paradox" where further increases in reliability will make the ratings less valid. Finally, the assumption of random error in making diagnoses does not always hold, so that statistical analyses must be adjusted accordingly. New statistical methods are needed to index only false-positive or false-negative rates in order to quantify the error that will reduce some validity coefficients.

Bipolar Disorder

Validity of occupational histories obtained by interview with female workers.

This study measured the validity of work histories obtained by interview with 84 female workers and examined specific factors which influence such validity. This is the first validation of work histories collected by interviews with women. The validity of each interview was assessed over a period of 29 years, from 1955 to 1983. The information provided by the worker was compared annually to job information registered in public and union records. On the average, interviews yielded the correct information (either employer's name or nonworking year) for 81% of the person years of these subjects. However, there was a time effect; the average validity score for recent employment (1972-1983) was 89%, while that for employment in the more distant past (1955-1971) was 74%. Furthermore, workers who had fewer jobs, had longer durations of employment, and were non-French speaking had higher validity scores. Most of these findings are consistent with previous studies conducted among male respondents.

Employment

Concurrent validity of two language screening tests.

The importance of ascertaining the validity of clinical instruments used to make decisions about individuals is discussed and the need for additional validation studies is emphasized. Steps that can be taken to confirm the validity for a particular application, setting, or population are described. As an example, the concurrent validity of two language screening instruments, the Fluharty Preschool Screening Test and the Northwestern Syntax Screening Test, and their subtests was examined. Decisions from these screening tests and subtests were compared to a validity criterion of passing or failing the Sequenced Inventory of Communication Development for 182 white middle-class children, ages 36-47 months. The results showed that the screening tests differed in their validity, depending upon the content of the test and each subtest. The consequences of using either screening test are explored, to illustrate how the outcomes of such studies should be interpreted.

Child, Preschool

Development and validation of an instrument to measure satisfaction of participants at breast screening programmes.

A reliable and valid questionnaire has been developed to measure the satisfaction of participants with service offered at mammography screening programmes. The questionnaire measures five specific aspects: convenience and accessibility, staffs' interpersonal skills, information transfer between staff and client, physical surroundings and perceived technical competence of staff. A general satisfaction dimension was also included. Systematic procedures were followed to ensure that the initial pool of items met the criteria for satisfactory content validity. These procedures included extensive literature review and interviews with participants and service providers. Discriminant validity was assessed by a modified Q-sort procedure, where eight expert judges sorted items into relevant dimensions. The sample for other validity and reliability testing consisted of 584 women who were participants at a breast X-ray programme in Melbourne, Australia. Concurrent validity was demonstrated by considering the correlation of the sum of the subscale scores for each respondent with their score on the general subscale (r = 0.76; P less than 0.001). Multiple regression was used to provide further evidence for the discriminant validity of the proposed subscales and support for the multidimensional conceptualism of satisfaction. Scores on the general satisfaction subscale were used as an outcome variable and other subscale scores were predictor variables. All subscale scores significantly contributed to the prediction of satisfaction, over and above that of other subscales (R2 = 0.59). This indicates that these subscales are measuring distinct dimensions of satisfaction. Cronbach's alpha of each subscale was over 0.50, indicating that the subscales are reliable. The instrument is a potentially useful tool for assessing the quality of care at mammographic screening services and could be used routinely by such services to monitor satisfaction.

Australia