PubMed HealthSearch

SEARCH · PubMed Health

Results for “validity”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Clinical Variable-Based Machine Learning for Predicting Early mCRPC Using Exclusively Clinical Variables: Development and Multicenter External Validation.

BACKGROUND AND OBJECTIVE: Metastatic hormone-sensitive prostate cancer (mHSPC) exhibits heterogeneous progression patterns, with early progression to metastatic castration-resistant prostate cancer (mCRPC) within 12 months indicating aggressive tumor biology and poor prognosis. Current risk stratification tools (CHAARTED, LATITUDE) offer limited individualized prediction. Machine learning approaches are increasingly applied to predict prostate cancer progression, but most models show modest performance (AUC 0.68-0.72), limited external validation, or require genomic variables unavailable in routine practice. This study aimed to develop and externally validate a novel RINH algorithm for predicting early mCRPC progression (≤ 12 months) using exclusively clinical variables, positioning it as a superior alternative to conventional ML classifiers. METHODS: This multicenter study enrolled 412 patients with de novo mHSPC from seven Spanish academic centers using mixed retrospective-prospective data collection. Twenty clinical variables were recorded, including demographics, PSA, ISUP grade, metastatic localization, CHAARTED/LATITUDE classifications, and treatment modalities. Following RINH-based outlier exclusion (55 patients), 357 patients (29 with early progression, 8.1%) were used to train six ML algorithms: RINH, Logistic Regression, Linear Discriminant, Support Vector Machine, Random Forest, and Subspace Discriminant. A two-tiered validation strategy integrated stratified fivefold cross-validation across all centers and formal external validation using center 1 (n = 121, 19 events) for training and centers 2-7 (n = 207, 10 events) for independent testing. Performance metrics included AUC, sensitivity, specificity, accuracy, and F1-score. KEY FINDINGS AND LIMITATIONS: Artificial intelligence and machine learning (ML) are transforming oncology, promising personalized risk stratification beyond traditional clinical criteria. In metastatic hormone-sensitive prostate cancer (mHSPC), early progression to castration resistance (mCRPC) within 12 months signals aggressive biology and poor prognosis, yet current tools (CHAARTED, LATITUDE) offer limited individualized prediction. Multiple ML models have been proposed with variable success: most achieve modest performance (AUC 0.68-0.72), lack robust external validation, or rely on genomic variables inaccessible in routine practice. We propose a novel approach using the Rivality Index Neighborhood (RINH) algorithm, demonstrating superior predictive capacity in an initial multicenter validation with exclusively clinical variables. This study provides rigorous multicenter external validation, advancing toward implementable precision oncology tools. CONCLUSIONS AND CLINICAL IMPLICATIONS: The RINH algorithm achieves superior predictive performance for early mCRPC progression using exclusively clinical variables, representing a significant advance toward implementable risk stratification. However, low reliability scores in external validation underscore that excellent performance metrics alone do not guarantee stability. Before clinical deployment, validation in substantially larger cohorts with higher progression events is essential. If validated, this model could enable personalized, risk-adapted therapeutic strategies, refining patient selection for treatment intensification or de-escalation.

Humans

Approaches for assessing the validity of a functional observational battery.

As neurobehavioral assessments during the preliminary stages of chemical testing are more widely undertaken, it is critical that the screening procedures utilized be valid indicators of neurobehavioral function and that they be sensitive, specific, and reliable. Efforts in this laboratory have been directed towards assessing these features in the use of a functional observational battery (FOB). For the purpose of assessing validity, we have examined FOB data which addresses the issues of criterion, predictive, concurrent, and construct validities. The FOB appears to be valid for detecting chemical-induced neurological dysfunction in rats, i.e., shows a good degree of criterion validity. Furthermore, in many instances the effects observed with the FOB may be predictive of symptomatology in humans. When comparisons can be made between effects detected with the FOB and other methods of measuring neurotoxicity (e.g., neuropathology), concurrent validity can also be established. To assess construct validity, effects of neurotoxicants can be classified into functional domains which are described by various measures in the FOB. Approaches for assessing the validity of the test method thus include answering specific research questions directed at assessing criterion, predictive, concurrent, and construct validity. Available data indicate that, in these aspects, the FOB is a valid screening method for the detection of neurotoxicity.

Animals

Reliability, Device Agreement and Validity of Load-Velocity Profiles: A Systematic Review with Meta-analysis.

BACKGROUND: For a valid one-repetition maximum (1RM) prediction via load-velocity (LV) relationships, high reliability and accuracy must be assumed. OBJECTIVE: Since individual study results indicate ambivalent prediction, this systematic review and meta-analysis was designed to provide a updated and comprehensive overview, extending knowledge about the validity and reliability of commercially available velocity sensors in Part I and the validity and reliability of velocity-based 1RM prediction models in Part II. METHODS: A systematic literature search was conducted in PubMed/MEDLINE, Web of Science, and Scopus. Validity and/or reliability studies or velocity-based 1RM prediction evaluations were included. Methodological quality was assessed using adapted COSMIN. The analysis was performed for intraclass correlation coefficient (ICC), Lin's concordance correlation coefficient (CCC), and Pearson's correlation coefficient (r). The review was preregistered in PROSPERO (CRD42025634595). RESULTS: Sixty-three studies were included for sensor validity and reliability and 38 for 1RM prediction models. Part I: Velocity sensors demonstrated good-to-excellent pooled validity and device agreement (ICC = 0.91-0.92 [0.83-0.97]; k = 55 and 439, respectively); intra- and inter-day reliability were classified as good to excellent with ICC = 0.90-0.91 [0.85-0.95] (k = 228 and 608, respectively), with sensor technology moderating the results. However, substantial heterogeneity and wide ranges of study-level estimates indicated considerable variability across moderators, linear position transducer (LPT) generally showing more consistent performance than inertial measurement units (IMU). Part II: Velocity-based 1RM prediction showed ICCs = 0.90 [0.83-0.94] (k = 124) and ICC = 0.91 [0.72-0.98] (k = 9); for reliability and validity, respectively. DISCUSSION: Commercial velocity sensors generally provide high relative validity and reliability. Results varied depending on exercise complexity, intensity, sensor technology, and modeling approach. While velocity-based 1RM prediction demonstrated high average validity, large heterogeneity in lower body exercises significantly biased the results. Furthermore, the dearth of measurement error and agreement analyses prohibits final conclusions. CONCLUSION: Therefore, velocity-based monitoring and 1RM prediction require cautious interpretation, as sensor- and exercise-specific evidence remains limited.

Load–velocity relationship

Validity of single variables and composite indices for measuring disease activity in rheumatoid arthritis.

There is no agreement as to which variable best mirrors disease activity in rheumatoid arthritis (RA) and no studies have been performed on the validity of disease activity variables. In this study the validity of 10 commonly used single variables and three composite indices was tested. All patients participated in a large follow up study in two clinics. The patients (n = 233) had classical or definite RA and a disease duration of less than one year at entry. The mean follow up time was 30 months; the follow up frequency was once every four weeks; 6011 records were used in the analysis. The validation criteria included correlations with the other variables (correlational validity), with the physical disability (criterion validity I), and with the radiographically determined damage of hands and feet (construct validity). The judgment of a group of rheumatologists in clinical practice was also used as a model of criterion validity (II). In this comparison the disease activity score and Mallya index showed the best validity. The best single variable was the number of swollen joints. The validity of most single variables was poor and these variables were not suitable as single endpoint measures in clinical trials.

Arthritis, Rheumatoid

Methods for defining equity-stratifying variables: a systematic review of validation studies.

BACKGROUND AND OBJECTIVE: Disease burden is often disproportionally higher among those who are socially disadvantaged by factors defined in the PROGRESS-Plus framework (ie, Place of residence, Race/ethnicity/culture/language, Occupation, Gender/sex, Religion, Education, Socioeconomic status, and Social capital, with "Plus" covering features like age and disability). The accuracy and applicability of case definitions to identify these variables from administrative and clinical health data are unknown. We conducted a systematic review to explore how equity-stratifying variables, as categorized by the PROGRESS-Plus framework, have been defined and validated in epidemiologic studies using administrative health, population-level, or electronic health record (EHR) data. METHODS: Medline, EMBASE, CINAHL, Web of Science, and Google Scholar were searched from the inception of the databases to 2024 for validation studies of equity-stratifying variables in adults using administrative health datasets, health registries, or EHR data. Titles and abstracts, followed by relevant full-text articles, were screened in duplicate by two reviewers for eligibility. The data sources utilized, algorithms employed, and their associated performance measures were extracted and synthesized from included studies. Given substantial heterogeneity in study design, equity-stratifying variable definition, and performance metrics, meta-analysis was not possible. RESULTS: Of the 9099 unique citations screened, 188 full texts were reviewed and 116 were included in this review. Most studies were published between 2019 and 2024 (n = 64, 55%) and were validation studies of race/ethnicity definitions that used race/ethnicity codes or surname list algorithms (n = 66, 57%). No studies examined religion. Regarding the reported performance measure estimates, the race/ethnicity/culture/language equity-stratifying variables category had the largest variability across sensitivity, positive predictive value (PPV), and Cohen's Kappa. Occupation validation studies had the lowest variation in sensitivity and PPV. CONCLUSION: Despite an increasing number of publications reporting on the validation of equity-stratifying variables relevant to the PROGRESS-Plus framework, performance measures varied widely across studies. The significant heterogeneity in equity-stratifying variable definitions and methods used to validate them support the need for further rigorous validation of equity-stratifying variables in administrative and clinical health data. PLAIN LANGUAGE SUMMARY: Disease burden is often higher in people who experience financial hardships, lower level of education, discrimination due to race/ethnicity, and unstable housing. These social factors can be considered health equity factors and are important for understanding health inequalities. Health researchers often use large datasets, such as hospital or electronic health records (EHRs), to study these health equity factors. However, it is not clear how accurately these data sources capture information about people's social circumstances and how these factors are defined. In this study, we reviewed existing research to understand how health equity factors have been defined across health data sources and how accurate they are at measuring aspects of health equity and social disadvantage. Of the more than 9000 studies we identified, we included 116 that met our criteria for this systematic review. Most included studies focused on identifying race and ethnicity, often using codes or surname-based methods. We found that the accuracy of these methods varied widely across studies, meaning results may not always be reliable or comparable. Overall, our findings show that there are inconsistencies in how social factors are defined and measured in health data. This makes it difficult to fully understand and address health inequalities using routinely collected health data. More work is needed to develop and validate better quality and more consistent methods for capturing these important social factors.

Humans

Kinesthetic aftereffect and personality: a case study of issues involved in construct validation.

Kinesthetic Aftereffect (KAE), once a promising personality index, has been abandoned by many investigators because of poor retest reliability and intermittent validity. In challenging this current consensus, we argue that (a) first-session KAE is valid; (b) poor retest reliability simply reflects later-session bias; (c) hence, multisession studies should not be used to assess validity without taking this bias into account. Those recent studies which failed to support KAE validity were each multisession in design. If our bias contention is correct, these studies should be ignored, and the claim of intermittent validity is thus rebutted. Reanalysis of the most recent major multisession, nonsupportive validity study indicates (a) Session 1 validity, (b) later-session bias, and (c) later-session valdiity when multisession scores are combined to avoid bias. Thus, KAE validly measures personality.

Humans

Self-reports by alcohol and drug abuse inpatients: factors affecting reliability and validity.

The reliability and validity of self-report data regarding substance abuse has often been questioned. To determine how best to enhance the veracity of self-report, three factors which might affect self-report veracity were examined: alcohol status at time of interview; level of cognitive functioning; and method of self-report data collection. Subjects were 234 admissions to an inpatient substance abuse treatment unit. Self-report data were collected via both personal interview on the day of admission and and questionnaire within the first week of stay. Self-reports concerned use of alcohol, cocaine, and marijuana in the days preceding admission. Test-retest reliability for the questionnaire data produced reliability coefficients of 0.88, 0.91, and 0.88, for alcohol, cocaine, and marijuana, respectively. Variation in inter-test interval had virtually no effect upon reliability coefficients. Interview data were compared to toxicologic analyses of blood and urine samples collected on admission. Overall, this comparison showed self-reports to be valid, with a 97% agreement between verbal report and laboratory data for alcohol, 93% for cocaine, and 84% for marijuana. The comparison of interview data with questionnaire responses also showed self-reports to be valid: 90% agreement for alcohol, 93% for cocaine, and 81% for marijuana. Level of cognitive function did not influence the validity of self-reports for any of the three substances. Recent consumption of alcohol also had no statistically significant effect on the validity of self-reported marijuana use, regardless of the operational form of validity tested. However, BAC-negative subjects produced a significantly greater validity coefficient for self-reported cocaine use (kappa = 0.87) than did BAC-positive patients (kappa = 0.43), when interview data were compared with toxicologic measures. A similar finding was not uncovered when interview and questionnaire data were compared. An interaction between admission alcohol status and cognitive function was uncovered for cocaine self-reports when interview data was compared with toxicologic measures. The rate of agreement for alcohol-negative subjects is quite high for both cognitively impaired and unimpaired subjects (M = 93% and M = 94%, respectively) as well as for alcohol-positive, cognitively unimpaired subjects (M = 94%), but not for alcohol-positive, cognitively impaired subjects (M = 67%). Results are discussed in terms of threats to the validity of self-report and strategies for the optimization of response accuracy.

Adult

Validity of information concerning the use of dental services obtained in interviews.

Two hundred and fifty-two persons out of a population of 358 were interviewed concerning their use of dental services. The validity of the information was tested by comparing the answers from each respondent with the contents of his/her dental treatment record. Replies to a question about the time interval since the last dental visit showed a high degree of validity. The validity of information concerning the type of treatment received at the last course of dental visits showed high validity for a single treatment and low validity when the treatment services were mixed. Responses about the regularity of treatment attendance demonstrated decreasing degree of validity with increasing number of dental visits during the last 5 years. The demographic and socioeconomic characteristics of the respondents showed little relation to the validity of their answers. However, the degree of validity decreased with increasing number of teeth.

Adult

'Truthsets' for clinical validation of large-scale functional assays: Practice recommendations from Cancer Variant Interpretation Group UK (CanVIG-UK).

BACKGROUND: Large-scale functional assays, including multiplex assays of variant effect, have substantial potential to resolve variants of uncertain significance (VUS), particularly for rare missense variants where clinical and population evidence are limited. The ClinGen assay-level clinical validation framework described by Brnich et al provided baseline guidance for the use of functional data for variant classification. However, clear consensus regarding construction of variant 'truthsets' by which to clinically validate functional data remains lacking. METHODS: CanVIG-UK developed consensus recommendations for truthset construction through an iterative national consultation process involving the CanVIG Steering Advisory Group (CStAG), wider CanVIG-UK membership, and engagement with international functional genomics experts. Consultation was based on previous analyses of 2,120 truthset constructions examining the impact of truthset composition on evidence point allocation within the ClinGen assay-level clinical validation framework. RESULTS: Across several consultations, CanVIG-UK established nine guiding principles and seven best-practice recommendations for assay-level clinical validation, using the assumed context of an assay for a cancer susceptibility gene where loss-of-function is the mechanism of pathogenicity. The principal recommendation stipulates, where assays are intended for use in interpretation of largely missense variants, the truthset used to validate should comprise only missense variants. Rather than mixtures of different variant types which may serve to over-estimate assay performance. Additional recommendations support option for relaxation of truthset stringency to improve power, augmentation of benign missense truthsets with systematically derived 'proxy-clinical' benign variants, independent clinical validation separate from assayist-defined validation, and careful evaluation of missense score distributions against that of protein-truncating and synonymous variants. Guidance is also provided for scenarios with limited pathogenic truthset availability and for assays reporting multiple deleterious zones or readouts. CONCLUSIONS: The CanVIG-UK principles and recommendations for truthset construction upon the ClinGen assay-level clinical validation framework, while aiming to form a baseline for future discussion regarding other functional and disease contexts and helping to address the gap between publication of new data and routine clinical implementation.

Journal Article

The preparation and validation of stock cultures of mammalian cells.

The utilization of continuous cell substrates is now widely accepted for the production of biologics. As part of the evaluation and licensing process of these products, the regulatory agencies are requiring extensive validation of the production system. That production system begins with the validation or qualification of a defined cell bank. This includes the validation of the stability cells at both the genetic and biochemical levels during cell culture production. Process validation plus cell bank validation provide the necessary assurance that the final product will be free of contaminating viruses and other adventitious agents. A combination of cell bank characterization and product characterization (peptide mapping or amino acid sequencing, or a combination thereof) will also demonstrate the stability of the production process. The validation of a cell bank for adventitious agents and cell line stability will not, in itself, ensure that the product is sterile and stable. However, cell bank validation is critical for demonstrating the safety of a product when combined with both process validation and end-product testing.

Animals

Industrial perspective on validation of tangential flow filtration in biopharmaceutical applications. Technical Report No. 15. Parenteral Drug Association. Biotechnology Task Force on Purification and Scale-up.

Validation of tangential flow filtration is required to ensure the process delivers a product of consistent quality, safety, and efficacy. A thorough and sound validation program not only satisfies regulatory requirements, but also provides a valuable source of information which facilitates development of future processes, training of production personnel, and trouble shooting for the validated process. Validation of TFF shares many common elements with validation of other traditional operations and equipment. Existing personnel and procedures should be readily adapted to execute the TFF validation protocols. IQ's and OQ's will most likely follow familiar formats. In performance qualification, key areas needing attention include: assessment of compatibles, testing of parameters affecting membrane retention and selectivity, cleaning, sanitization, and membrane lifetime. Finally, the hallmark of a sound validation program is the quality of its scientific approach and its congruence with the definition of validation contained in the 1987 guidelines (6).

Filtration

[Validity of the Hoppe List. An empirical study of patients with backache].

Construct validity, criterion-oriented validity, and differential validity were assessed to determine the validity of the Hoppe-Liste (HL). Construct validity was ascertained by the multitrait-multimethod analysis suggested by Campbell & Fiske. Data regarding both convergent and discriminant validity show that the HL measures the construct pain. A comparison of the HL with the Revised Multidimensional Pain Scale and the Visual Analogue Scale suggests the criterion-oriented validity of the HL. Group-specific scores of the HL scales and of correlation patterns between scales tentatively support differential validity. It is concluded that the HL seems to be a suitable instrument for measuring the intensity and quality of pain.

Back Pain

Validation of a Turkish Translation of the Stress in Emergency Healthcare Professionals: The Stress Factors and Manifestations Scale.

AIM: The primary duties of emergency healthcare professionals (EHPs) are to provide emergency patient care to acutely ill and injured individuals. Due to the nature of their work, EHPs operate under constant stress, often requiring rapid decision-making, swift action, and the delivery of necessary medical care in life-or-death situations, sometimes under inadequately safe conditions. Therefore, the aim of this study is to determine the validity and reliability of the Emergency Healthcare Professional Stress Factors and Symptoms (SEHP:SFMS) Scale in Turkish for identifying stress factors and symptoms in emergency medical care professionals providing emergency patient care services. DESIGN: A methodological study design was used in this study. METHODS: The study was conducted with the participation of 211 EHPs from employees working in emergency care institutions affiliated with the Muğla Provincial Health Directorate between November 2023 and June 2024. Data were collected via a face-to-face survey. Data were analysed using Lawshe content validity ratio, Kaiser-Meyer-Olkin coefficient, Bartlett test, exploratory factor analysis, principal component analysis, Varimax factor rotation method, confirmatory factor analysis, Cronbach's α internal consistency coefficient, convergent validity, discriminant validity, test-retest, and Spearman correlation coefficient tests. RESULTS: The linguistic translation and cultural adaptation of the SEHP:SFMS showed strong performance. The scope validity index of the scale is 0.83. The item-total correlation values of the scale were found to be between 0.486 and 0.794, and the factor loadings were between 0.474 and 0.816. Confirmatory factor analysis fit indices: χ2 = 248.727; df = 101; n = 211; p = 0.000; χ2/df = 2.463; RMSEA = 0.083; CFI = 0.914, SRMR = 0.052, which was found to be compatible and acceptable with the proposed 3-factor model. The Cronbach's α reliability coefficient of the scale was 0.931, and the total variance was 61.97%. CONCLUSIONS: SEHP:SFMS is a valid and reliable tool to assess stress factors and symptoms of Turkish emergency healthcare professionals. Its use improves the quality of emergency care. PATIENT OR PUBLIC CONTRIBUTION: These study findings have been used to create a tool with Turkish validity and reliability that allows for the examination of stress factors among healthcare professionals working in emergency and critical services. Identifying and reducing stress factors among healthcare professionals is crucial for the delivery of quality healthcare services. It can also be used to develop targeted interventions and ongoing strategies to facilitate improved clinical supervision and mentoring. IMPLICATION FOR NURSING PRACTICE: Nurses in emergency departments, which are among the most stressful, dynamic, intense, life-saving, and critical environments in healthcare institutions, and where life-saving treatment is administered, are at high risk of experiencing psychological trauma. Trauma experienced in the work environment is a significant problem for nursing. The consequences of trauma negatively affect nurses and institutions. Studies show that post-traumatic stress, anxiety, depression, and burnout are commonly observed in emergency department nurses. In this sense, understanding the stress and stress factors experienced by nurses can guide future interventions. The results of this study are considered important in making visible the stress and stress factors experienced by nurses in the emergency department, and also in guiding managers and nurses working in this field in terms of preventive and protective measures.

Humans

Reliability and descriptive validity of PSE syndromes.

Despite extensive research use of the Present State Examination (PSE), the validity of classification based on PSE data has not been studied extensively. We have examined a consecutive series of functional psychiatric admissions using the PSE and systematically gathered clinical and demographic data in order to study not only the reliability but the descriptive (construct) validity of classification based on PSE data. We have found that the PSE can be used in a psychiatric hospital to reliably describe and classify schizophrenic and affective syndromes with considerable descriptive validity in terms of clinical and demographic variables. We believe that this type of validity is an important step in establishing validity of clinical disease entities. The interrelationship among different kinds of validity (descriptive, concurrent, predictive) might provide clinical disease concepts with more definitive validity.

Adult

Analytical methods validation: bioavailability, bioequivalence and pharmacokinetic studies. Conference report.

This is a summary report of the conference on Analytical Methods Validation: Bioavailability, Bioequivalence and Pharmacokinetic Studies. The conference was held from December 3 to 5, 1990 in the Washington, DC area and was sponsored by the American Association of Pharmaceutical Scientists, US Food and Drug Administration, Federation International Pharmaceutique, Health Protection Branch (Canada) and Association of Official Analytical Chemists. The purpose of the report is to represent our assessment of the major agreements and issues discussed at the conference. The report is also intended to provide guiding principles for validation of analytical methods employed in bioavailability, bioequivalence and pharmacokinetic studies in man and animals. The objectives of the conference were: 1. To reach a consensus on what should be required in analytical methods validation and the procedures to establish validation; 2. To determine processes of application of the validation procedures in the bioavailability, bioequivalence and pharmacokinetic studies; 3. To develop a report on analytical methods validation (which may be referred to in developing future formal guidelines). Acceptable standards for documenting and validating analytical methods with regard to processes, parameters or data treatments were discussed because of their importance in assessment of pharmacokinetic, bioavailability and bioequivalence studies. Other topics which were considered essential in the conduct of pharmacokinetic studies or in establishing bioequivalency criteria, including measurement of drug metabolites and stereoselective determinations, were also deliberated.

Biological Availability

Issues of validity in the Diagnostic Interview Schedule.

The Diagnostic Interview Schedule, the chief instrument in contemporary studies in psychiatric epidemiology, enhances the reliability of psychiatric diagnosis and enables lay interviewers to closely reproduce psychiatric interviews. However, despite frequent references in the literature to the validity of the Diagnostic Interview Schedule, most studies fundamentally represent variations of reliability paradigms to the neglect of criterion-related validity. Mistaken assertions of validity persist in the psychometric language used to describe the Diagnostic Interview Schedule. This article examines the basis for claims and counterclaims of validity in accordance with standard psychometric definition, and identifies sources of erroneous reasoning in attempts to infer validity from reliability. The article presents a general framework organizing the process of diagnostic validation and discusses strategies for research seeking to validate psychiatric diagnoses achieved through the Diagnostic Interview Schedule.

Humans

Automating candidate gene prioritization with large language models: from naive scoring to literature-grounded validation.

MOTIVATION: Identifying promising therapeutic targets from thousands of genes in transcriptomic studies remains a major bottleneck in biomedical research. While large language models (LLMs) show potential for gene prioritization, they suffer from hallucination and lack systematic validation against expert knowledge. RESULTS: The framework identified 609 sepsis-relevant genes with >94% filtering efficiency, demonstrating strong enrichment for inflammatory pathways including TNF-α signaling, complement activation, and interferon responses. Literature validation yielded 30 ultra-high confidence therapeutic candidates, including both established sepsis genes (IL10, TREM1, S100A9, NLRP3) and novel targets warranting investigation. Benchmark validation against expert-curated databases achieved 71.2% recall, with systematic correlation between computational confidence and evidence quality. The final candidate set balanced discovery (11 novel genes) with validation (19 known genes), maintaining biological coherence throughout the filtering process. This framework demonstrates that rigorous methodology can transform unreliable LLM outputs into systematically validated biological insights. By combining computational efficiency with literature grounding, the approach provides a practical tool for prioritizing experimental validation efforts. The modular design enables adaptation to other diseases through knowledge base substitution, offering a systematic approach to literature-guided biomarker discovery. AVAILABILITY AND IMPLEMENTATION: We developed a two-stage computational framework that combines LLM-based screening with literature validation for systematic gene prioritization. Starting with 10 824 genes from the BloodGen3 repertoire, we applied multi-criteria evaluation for sepsis relevance, followed by retrieval-augmented generation using 6346 curated sepsis publications. A novel faithfulness evaluation system verified that LLM predictions aligned with retrieved literature evidence. Source code and implementation details are available at https://github.com/taushifkhan/llm-geneprioritization-framework, vector database at https://doi.org/10.5281/zenodo.15802241, and Interactive demonstration at https://llm-geneprioritization.streamlit.app/.

Humans

Development and content validity testing of a comprehensive classification of diagnoses for pediatric nurse practitioners.

Pediatric nurse practitioners (PNPs) need an integrated, comprehensive classification that includes nursing, disease, and developmental diagnoses to effectively describe their practice. No such classification exists. Further, methodologic studies to help evaluate the content validity of any nursing taxonomy are unavailable. A conceptual framework was derived. Then 178 diagnoses from the North American Nursing Diagnosis Association (NANDA) 1986 list, selected diagnoses from the International Classification of Diseases, the Diagnostic and Statistical Manual, Third Revision, and others were selected. This framework identified and listed, with definitions, three domains of diagnoses: Developmental Problems, Diseases, and Daily Living Problems. The diagnoses were ranked using a 4-point scale (4 = highly related to 1 = not related) and were placed into the three domains. The rating scale was assigned by a panel of eight expert pediatric nurses. Diagnoses that were assigned to the Daily Living Problems domain were then sorted into the 11 Functional Health patterns described by Gordon (1987). Reliability was measured using proportions of agreement and Kappas. Content validity of the groups created was measured using indices of content validity and average congruency percentages. The experts used a new method to sort the diagnoses in a new way that decreased overlaps among the domains. The Developmental and Disease domains were judged reliable and valid. The Daily Living domain of nursing diagnoses showed marginally acceptable validity with acceptable reliability. Six Functional Health Patterns were judged reliable and valid, mixed results were determined for four categories, and the Coping/Stress Tolerance category was judged reliable but not valid using either test. There were considerable differences between the panel's, Gordon's (1987), and NANDA's clustering of NANDA diagnoses. This study defines the diagnostic practice of nurses from a holistic, patient-centered perspective. It is the first study to use quantitative methods to test a diagnostic classification system for nursing. The classification model could also be adapted for other nurse specialties.

Humans