PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “External validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2×2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I²=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans↗

Who enrolls in prevention trials? Discordance in perception of risk by professionals and participants.

Internal and external validity problems permeate all intervention studies but are accentuated in primary preventive intervention research, particularly when studies target or recruit individuals based on their risk for psychopathology. Since many people who are at risk do not yet experience distress, they may not perceive the need for intervention. Recruitment tactics based on explaining extent of risk are unlikely to be persuasive and may have negative consequences. If respondents are not motivated to participate, a small or biased subset of the target population will participate in the intervention. Bias is of special concern when those enrolled represent only part of the continuum of risk. Selective enrollment may compromise both internal validity (the interpretation of the research results) and external validity (the generalizability of the findings) of intervention trials in primary prevention. This article discusses the effects of partial enrollment and the resultant bias. It suggests several strategies for increasing the enrollment of the target population and examines some of their ethical ramifications. It also stresses the importance of collecting systematic data documenting how the participants in the intervention differ from the target group as a whole.

Bias↗

Short form of a situational temptation scale for heavy, episodic drinking.

PURPOSE: A short form for situational temptations to drink scale was developed from an original 21-item inventory by Migneault. METHODS: The form measured four hypothesized subscales of temptations on a sample of 348 college drinkers (66% female). Peer pressure, social anxiety, negative affect, and positive/social situations subscales were replicated and reduced. RESULTS: Strong empirical support was found for a hierarchical model, indicating that the four subscales can be summed to provide a global measure of situational temptations. Confirmatory factor results, internal and external validity, and high correlations with the original measures indicate that the short form was as psychometrically valid as the original measure. IMPLICATIONS: Measures of external validity demonstrated the applicability of this measure to heavy drinking prevention programs.

Adult↗

Changes in obsessive/compulsive patients as measured by the Leyton Inventory before and after treatment with clomipramine.

The Leyton Obsessional Inventory has been found to be a useful measure in assessing patients before and after treatment with clomipramine. Mean scores for symptoms and interference altered significantly during the course of treatment. The Leyton Obsessional Inventory, however, lacks external validation owing to the absence of some valid alternative quantification. In the absence of such external validation it seems justifiable to use the mean Leyton score diagnostically but not as a sole indication of severity or response to treatment.

Clinical Trials as Topic↗

Model validation for external doses due to environmental contaminations by the Chernobyl accident.

The objective of the present paper is to validate the deterministic JSP5 model for external exposures to population groups living in the areas contaminated with radionuclides after the Chernobyl accident. For this purpose inhabitants of contaminated areas wore TL-dosimeters for about 1 mo in the spring/summer periods of the years 1989 to 1994. External doses due to the Chernobyl accident were determined from the dosimeter readings by subtracting the natural background. 2,342 results for rural inhabitants and 420 results for inhabitants of the town Novozybkov passed reliability checks. These data show that the average dose in inhabitants of a rural settlement predicted by the model is in the range 0.69-1.55 of the measured values with a confidence level of 95%. Differences are attributed to settlement specific location factors, which are supported by the very good agreement of model and measurements in Novozybkov. In this case location factors of the model were obtained from Novozybkov directly.

Adult↗

Is the ACLS score a valid prediction rule for survival after cardiac arrest?

UNLABELLED: The ACLS (advanced cardiac life support) Score was previously developed to predict survival from out-of-hospital cardiac arrest. Whether the arrest was witnessed, initial cardiac rhythm, performance of bystander cardiopulmonary resuscitation (CPR), and the response time of the paramedic unit were determined to be predictive of survival. However, the ACLS Score has not been validated in other emergency medical services systems. OBJECTIVES: The purpose of this study was to externally validate the ACLS Score in one patient population. METHODS: This was a retrospective cohort study performed at an urban county teaching hospital. The study population consisted of consecutive adult patients treated for out-of-hospital, nontraumatic cardiac arrest, and transported to the authors' institution between November 1, 1994, and September 30, 2001. Patient records for all cardiac arrests during the study period were reviewed. Study variables included witnessed arrest, initial arrest rhythm, bystander CPR, paramedic response time, and survival to hospital discharge. Predicted probability of survival to hospital discharge was calculated for each patient using the ACLS Score. The overall predicted and observed survival rates were compared using Flora's Z score. The Hosmer-Lemeshow test was used to evaluate the model's goodness-of-fit over a range of survival probabilities. RESULTS: Of 754 cardiac arrest patients enrolled in the study period, 575 (76%) patients had documentation that allowed scoring using the ACLS Score. Twenty-five (4%) patients survived to hospital discharge. The predicted number of survivors based on the ACLS Score was 104 (18%), yielding a Flora's Z statistic of -4.46 (p < 0.0001). After categorizing predicted survival probabilities into four categories, the resulting Hosmer-Lemeshow statistic was 210 (p << 10(-6)). Both goodness-of-fit statistics demonstrated extremely poor fit of the model. A receiver operating characteristic (ROC) curve was created, yielding an area under the ROC curve of 0.33 (95% CI = 0.19 to 0.47), signifying extremely poor discrimination. CONCLUSIONS: The previously published ACLS Score was not valid when applied to an external cohort of out-of-hospital cardiac arrest patients. An externally valid model is needed to predict survival to hospital discharge following out-of-hospital cardiac arrest.

Advanced Cardiac Life Support↗

Machine Learning-Driven Prediction of Coronary Artery Disease Risk Based on UK Biobank Plasma Proteomics.

BACKGROUND: Coronary artery disease (CAD) is a leading global cause of mortality, yet the predictive accuracy of conventional risk models is limited. Here, we integrate conventional risk factors, polygenic risk scores, and large-scale proteomics to develop a unified model for enhanced CAD risk prediction. METHODS: Using data from UK Biobank, participants with plasma proteomics and genetic risk data were included after excluding prevalent CAD. Participants from England were split into training (n=32&#x2009;330) and internal validation (n=13&#x2009;857) sets, and Scotland/Wales participants formed an external validation set (n=5775). Incident CAD was ascertained from linked health records. A 202-protein proteomic risk score was derived by least absolute shrinkage and selection operator Cox regression, and CatBoost models were trained using conventional risk factors alone and with incremental addition of polygenic risk scores and protein proteomic risk scores; Shapley Additive Explanations-guided forward selection identified a compact protein panel. RESULTS: Across cohorts, the median age was 58&#x2009;years and &#x223c;45% were men. Protein proteomic risk score was dose-dependently associated with CAD risk. Compared with conventional risk factors alone, integrating polygenic risk scores and protein proteomic risk scores improved discrimination, with the area under the curve increasing from 0.750 (95% CI, 0.732-0.767) to 0.789 (95% CI, 0.772-0.805) in internal validation and from 0.717 (95% CI, 0.683-0.750) to 0.762 (95% CI, 0.732-0.791) in external validation. A 9-protein panel (GDF15 [growth differentiation factor 15], MMP12 [matrix metalloproteinase 12], NPPB [natriuretic peptide B], PGF [placental growth factor], REN [renin], ADGRG2 [adhesion G-protein coupled receptor], ACE2 [angiotensin-converting enzyme 2], CDCP1 [CUB domain-containing protein 1], CXCL17 [C-X-C motif chemokine ligand 17)]) captured most proteomic predictive information. CONCLUSIONS: Our findings demonstrate that integrating conventional risk factors, polygenic risk scores, and proteomic data improves CAD risk prediction. This study highlights the utility of proteomics in precision cardiovascular medicine and simplified risk stratification tools.

Humans↗

Generalizing disease management program results: how to get from here to there.

For a disease management (DM) program, the ability to generalize results from the intervention group to the population, to other populations, or to other diseases is as important as demonstrating internal validity. This article provides an overview of the threats to external validity of DM programs, and offers methods to improve the capability for generalizing results obtained through the program. The external validity of DM programs must be evaluated even before program selection and implementation are begun with a prospective new client. Any fundamental differences in characteristics between individuals in an established DM program and in a new population/environment may limit the ability to generalize.

Disease Management↗

Machine learning-enabled multi-omics discovery of prognostic biomarkers and signaling targets in pancreatic cancer.

Pancreatic ductal adenocarcinoma (PDAC) remains difficult to subtype using single omics layers. We conducted an exploratory investigation integrating reverse-phase protein array (RPPA) and DNA methylation data from the cancer genome atlas (TCGA)- pancreatic adenocarcinoma (PAAD) to assess the feasibility of multi-omics subtyping, alongside a supervised machine learning analysis of a small gene expression omnibus (GEO) transcriptomic cohort (n&#x202f;=&#x202f;26) to identify candidate diagnostic genes. RPPA-based K-means clustering suggested a weak, possible two-subtype structure (silhouette &#x2248; 0.16) that remained unassociated with overall survival (log-rank p&#x202f;=&#x202f;0.113) and lacked independent prognostic value. An independently performed similarity network fusion (SNF) analysis integrating RPPA and methylation data showed low concordance with RPPA-derived subtypes (Adjusted Rand Index (ARI) =&#x202f;0.014), indicating limited convergence between molecular modalities. Supervised machine learning analysis of the GEO cohort using a fully nested leave-one-out cross-validation pipeline achieved a mean (area under the curve) AUC of 0.896 across four classifiers and identified four-fold-stable candidate genes (ESCO2, COL17A1, BCL2L14, and SOWAHB). However, this gene panel demonstrated limited external validity across two independent PDAC cohorts (log-rank p&#x202f;=&#x202f;0.438 for both GSE62452 and GSE28735), indicating limited generalizability despite robust internal performance. Collectively, these findings provide limited evidence for a robust, prognostically significant multi-omics subtype or a validated diagnostic gene signature; instead, this study serves as a hypothesis-generating resource and highlights the importance of rigorous cross-validation and independent external validation in small-sample transcriptomic biomarker discovery.

Humans↗

Pragmatic controlled clinical trials in primary care: the struggle between external and internal validity.

BACKGROUND: Controlled clinical trials of health care interventions are either explanatory or pragmatic. Explanatory trials test whether an intervention is efficacious; that is, whether it can have a beneficial effect in an ideal situation. Pragmatic trials measure effectiveness; they measure the degree of beneficial effect in real clinical practice. In pragmatic trials, a balance between external validity (generalizability of the results) and internal validity (reliability or accuracy of the results) needs to be achieved. The explanatory trial seeks to maximize the internal validity by assuring rigorous control of all variables other than the intervention. The pragmatic trial seeks to maximize external validity to ensure that the results can be generalized. However the danger of pragmatic trials is that internal validity may be overly compromised in the effort to ensure generalizability. We are conducting two pragmatic randomized controlled trials on interventions in the management of hypertension in primary care. We describe the design of the trials and the steps taken to deal with the competing demands of external and internal validity. DISCUSSION: External validity is maximized by having few exclusion criteria and by allowing flexibility in the interpretation of the intervention and in management decisions. Internal validity is maximized by decreasing contamination bias through cluster randomization, and decreasing observer and assessment bias, in these non-blinded trials, through baseline data collection prior to randomization, automating the outcomes assessment with 24 hour ambulatory blood pressure monitors, and blinding the data analysis. SUMMARY: Clinical trials conducted in community practices present investigators with difficult methodological choices related to maintaining a balance between internal validity (reliability of the results) and external validity (generalizability). The attempt to achieve methodological purity can result in clinically meaningless results, while attempting to achieve full generalizability can result in invalid and unreliable results. Achieving a creative tension between the two is crucial.

Blood Pressure↗

Data collection methods in prospective economic evaluations: how accurate are the results?

OBJECTIVES: Often in economic evaluations a division is made between those studies that have a high level of accuracy versus those that are easily generalized. This interstudy dichotomy is often translated into prospective, randomized controlled trials with high internal validity and observational and modeling studies with a high level of external validity. This article challenges this conventional view and examines intrastudy effects on validity. METHOD: A review and summary of the literature was conducted in order to assess the impact that data collection strategies will have on internal validity. Two scenario models were created in order to gain a preliminary understanding of the magnitude of the problem. RESULTS: Data collection strategies have an impact on the level of internal validity found in an economic evaluation. Comparisons of studies that are prospective in nature is misleading as data collection strategy can lead to different resource and cost estimates even when all other relevant factors are similar. It is possible to shift and improve the level of validity by combining different collection methods. CONCLUSIONS: Instead of viewing internal and external validity as polar opposites, validity should be considered in terms of a continuum within a particular study. The use of proxies to collect resource utilization estimates, the reliance on patient self-reported data, and the method of collecting this type of data all impact the validity of study results. National guidelines for the economic evaluation of agents and devices should consider this issue in more depth, and existing evidence rankings should be adapted to be more appropriate to pharmacoeconomic studies.

Journal Article↗

Quantification of communication processes, is it possible?

The PAS system (Problem-Analysis-Solution-system) is developed to quantify oral communication processes during counselling in pharmacy practice. The pharmacist translates the patient's drug-related questions into a P-code, the analysis of the question into an A-code and finally the given solution upon the question into a S-code. The PAS system has been developed for two goals. First, for the registation of drug-related questions from patients which gives the pharmacist insight in the most common issues addressed by patients. Second, it might help the pharmacist to structure the communication with the patient during the consultation. Forty-one pharmacists participated in the evaluation of the PAS system. The validation of the PAS system consisted of two phases: the external validation and the internal validation. Kappa values were calculated as a measure of agreement in the coding by the pharmacists. The kappa-value of the external validation for the P-, A- and S-codes for the total set of questions indicate a moderate to poor agreement. This means that pharmacists categorize drug-related questions from patients in a different way. Therefore we conclude that the PAS system is less reliable for research purpose. The kappa-value of the internal validation for the P-code varies from 0.42 to 0.91. For the A-code it varies from 0.07 to 0.35 and for the S-code from zero to 0.68. Internal reproducibility is good for P-code but not for the A-code and S-code. This implies that the pharmacist can use the P-codes for registration of patients' questions in his own pharmacy. Moreover, the usage of the PAS system during counselling in pharmacy practice can structure the consultation.

Communication↗

An observational examination of the literature in diagnostic anatomic pathology.

Original research published in the medical literature confronts the reader with three very basic and closely linked questions--are the authors' conclusions true in the contextual setting in which the work was performed (internally valid); if so, are the conclusions also applicable in other practice settings (externally valid); and, if the conclusions of the study are bona fide, do they represent an important contribution to medical practice or are they true-but-insignificant? Most publications attempt to convince readers that the researchers' conclusions are both internally valid and important, and occasionally papers also directly address external validity. Developing standardized methods to facilitate the prospective determination of research importance would be useful to both journals and their readers, but has proven difficult. In contrast, the evidence-based medicine (EBM) movement has had more success with understanding and codifying factors thought to promote research validity. Of the many variables that can influence research validity, research design is the one that has received the most attention. The present paper reviews the contributions of EBM to understanding research validity, looking for areas where EBM's body of knowledge is applicable to the anatomic pathology (AP) literature. As part of this project, the authors performed a pilot observational analysis of a representative sample of the current pertinent literature on diagnostic tissue pathology. The results of that review showed that most of the latter publications employ one of the four categories of "observational" research design that have been delineated by the EBM movement, and that the most common of these observational designs is a "cross-sectional" comparison. Pathologists do not presently use the "experimental" research designs so admired by advocates of EBM. Slightly > 50% of AP observational studies employed statistical evaluations to support their final conclusions. Comparison of the current AP literature with a selected group of papers published in 1977 shows a discernible change over that period that has affected not just technological procedures, but also research design and use of statistics. Although we feel that advocates of EBM deserve credit for bringing attention to the close link between research design and research validity, much of the EBM effort has centered on refining "experimental" methodology, and the complexities of observational research have often been treated in an inappropriately dismissive manner. For advocates of EBM, an observational study is what you are relegated to as a second choice when you are unable to do an experimental study. The latter viewpoint may be true for evaluating new chemotherapeutic agents, but is unacceptable to pathologists, whose research advances are currently completely dependent on well-conducted observational research. Rather than succumb to randomization envy and accept EBM's assertion that observational research is second best, the challenge to AP is to develop and adhere to standards for observational research that will allow our patients to benefit from the full potential of this time tested approach to developing valid insights into disease.

Anatomy↗

Treatment satisfaction of patients with lower urinary tract symptoms: randomised controlled trials vs. real life practice.

Randomised controlled trials (RCTs) are an important scientific tool to determine the efficacy and tolerability of a given treatment relative to placebo or other treatment forms. However, due to strict inclusion and exclusion criteria the patient populations in RCTs may not be fully representative for those routinely consulting the physician. Moreover, participation in a formal study puts physician and patient in a situation where they may react different than in real life. In contrast real life practice (RLP) studies cannot determine treatment efficacy or tolerability in absolute terms since they typically do not include a control group and are purely observational. On the other hand, they tend to be more representative for real treatment outcomes. Thus, RCTs have high internal but less external validity whereas RLP studies have less internal and greater external validity. Hence, RCTs and RLP studies should not be considered as mutually exclusive but rather as complementing each other. Specific advantages and disadvantages of RCTs and RLP studies will be discussed using published evidence for the treatment of lower urinary tract symptoms suggestive of benign prostatic obstruction with alpha1-adrenoceptor antagonists and other treatments.

Humans↗

A clinical and echocardiographic score for assigning risk of major events after dobutamine echocardiograms.

OBJECTIVES: We sought to develop and validate a risk score combining both clinical and dobutamine echocardiographic (DbE) features in 4890 patients who underwent DbE at three expert laboratories and were followed for death or myocardial infarction for up to five years. BACKGROUND: In contrast to exercise scores, no score exists to combine clinical, stress, and echocardiographic findings with DbE. METHODS: Dobutamine echocardiography was performed for evaluation of known or suspected coronary artery disease in 3156 patients at two sites in the U.S. After exclusion of patients with incomplete follow-up, 1456 DbEs were randomly selected to develop a multivariate model for prediction of events. After simplification of each model for clinical use, the models were internally validated in the remaining DbE patients in the same series and externally validated in 1733 patients in an independent series. RESULTS: The following score was derived from regression models in the modeling group (160 events): DbE risk = (age.0.02) + (heart failure + rate-pressure product <15000).0.4 + (ischemia + scar).0.6. The presence of each variable was scored as 1 and its absence scored as 0, except for age (continuous variable). Using cutoff values of 1.2 and 2.6, patients were classified into groups with five-year event-free survivals >95%, 75% to 95%, and <75%. Application of the score in the internal validation group (265 events) gave equivalent results, as did its application in the external validation group (494 events, C index = 0.72). CONCLUSIONS: A risk score based on clinical and echocardiographic data may be used to quantify the risk of events in patients undergoing DbE.

Cardiotonic Agents↗

Cost estimates for hospital inpatient care in Australia: evaluation of alternative sources.

OBJECTIVE: This paper presents a framework for evaluation of alternative sources of estimates of the costs of hospital inpatient care in Australia. It argues that the choice of costing methods depends on the decision-context and the sensitivity of the decision to estimation errors. METHOD: Five criteria are proposed for evaluation of sources of hospital cost data, with detailed consideration of the way estimates are derived in two computerised approaches which use accounting data. Three broad approaches to cost estimation are evaluated against these criteria. RESULTS: Choosing an estimation method entails an optimisation analysis for each decision context. 'Microcosting' techniques remains the most valid approach to cost estimation, but are costly and this may, in turn, limit the sample of patients or institutions. Protocol-based cost estimates vary widely in their validity, depending on source data, but there is little justification for continued use of crude per diem cost estimates in such protocols. When precision and resolution are important objectives, clinical costing approaches provide the most valid inpatient cost estimates at a reasonable data cost. When external validity is important, or where standardisation of hospital costs is desired, use of published national cost weights may be preferred. CONCLUSION: Both primary and secondary sources of cost data must withstand challenges to internal and external validity. The 'resolution' (or precision) of cost estimates and the relative costs of collection must also be considered. IMPLICATIONS: Studies using estimates of the costs of hospital care should defend the appropriateness of the costing approach and data source for the decision context.

Accounting↗

Treatment research at the crossroads: the scientific interface of clinical trials and effectiveness research.

OBJECTIVE: Policy and clinical management decisions depend on data on the health and cost impacts of psychiatric treatments under usual care, i.e., effectiveness. Clinical trials, however, provide information on treatment efficacy under best-practice conditions. An understanding of the design, analysis, and conventions of both efficacy and effectiveness studies can lead to research that better informs clinical and societal questions. METHOD: This paper contrasts the strengths and limitations of clinical trials and effectiveness studies for addressing policy and clinical decisions. These research approaches are assessed in terms of outcomes, treatments, service delivery context, implementation conventions, and validity. RESULTS: Clinical trials and effectiveness research share problems of internal and external validity despite more attention to internal validity in clinical trials (e.g., randomization, blinding, standardized protocols) and to external validity in effectiveness studies (e.g., community-based treatments, representative samples). CONCLUSIONS: To develop research at the interface of clinical trials and effectiveness studies, research goals must be redefined, and methods, such as cost-utility and econometric analyses, must be shared and developed. Development of hybrid designs that combine features of efficacy and effectiveness research will require separation of conventions such as frequency of follow-up, intensity of measurement, and sample size from the central scientific issues of aims and validity.

Clinical Protocols↗

Validation of psychoanalytic theories: towards a conceptualization of references.

The authors discuss criteria for the validation of psychoanalytic theories and develop a heuristic and normative model of the references needed for this. Their core question in this paper is: can psychoanalytic theories be validated exclusively from within psychoanalytic theory (internal validation), or are references to sources of knowledge other than psychoanalysis also necessary (external validation)? They discuss aspects of the classic truth criteria correspondence and coherence, both from the point of view of contemporary psychoanalysis and of contemporary philosophy of science. The authors present arguments for both external and internal validation. Internal validation has to deal with the problems of subjectivity of observations and circularity of reasoning, external validation with the problem of relevance. They recommend a critical attitude towards psychoanalytic theories, which, by carefully scrutinizing weak points and invalidating observations in the theories, reduces the risk of wishful thinking. The authors conclude by sketching a heuristic model of validation. This model combines correspondence and coherence with internal and external validation into a four-leaf model for references for the process of validating psychoanalytic theories.

Humans↗