PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

The reliability of three active motor tests used in painful shoulder disorders. Presentation of a method of general applicability for the analysis of reliability in the presence of pain.

This article deals with reliability aspects of standardized, active motor tests ("functional tests") when applied to patients with painful shoulder disorders. Motor performance was rated independently by the same two examiners in a standardized way in three different manoeuvres: the Hand in Neck, Hand in Back, and Pour out of a Pot tests. Pain experienced during these tests was rated by the patients on a verbal scale. A method of general applicability is presented for the analysis of reliability of standardized, active motor tests when applied to painful shoulder joint disorders. The importance of differential motivation is stressed, as is the importance of using reliability measures that are adapted to the specific purpose of a particular clinical investigation.

Adult↗

Development and reliability of a standard rating system for outcome measurement of foot and ankle disorders II: interclinician and intraclinician reliability and validity of the newly established standard rating scales and Japanese Orthopaedic Association rating scale.

BACKGROUND: This study evaluated the validity and inter- and intraclinician reliability of (1) the Japanese Society of Surgery of the Foot (JSSF) standard rating system for four sites [ankle-hindfoot (AH), midfoot (MF), hallux (HL), and lesser toe (LT)] and the rheumatoid arthritis (RA) foot and ankle scale and (2) the Japanese Orthopaedic Association's foot rating scale (JOA scale). METHODS: Clinicians from the same institute independently evaluated participating patients from their institute by two evaluations at a 1- to 4-week interval. Statistical evaluation was as follows. (1) The intraclass correlation coefficient (ICC) was calculated from data collected from at least two examinations of each patient by at least two evaluating clinicians (Data A). (2) Total scores for the two evaluations were determined from the distribution of differences in data between the two evaluations (Data B); each item was evaluated by determining Cohen's coefficient of agreement. (3) The relation between patient satisfaction and total score was investigated only for patients who underwent surgery (Data C). Spearman's rank correlation coefficient was obtained. RESULTS: Participants were 65 clinicians and 610 patients, including those with disorders of the AH (313), MF (47), HL (153), and LT (50) and those with RA (47). From Data A, the ICC was high for AH and HL by JSSF scales and for AH, MF, and LT by the JOA scale. From Data B, the coefficient showed high validity for both scales for AH, with almost no difference between the two scales; the validity for HL was higher with the JOA scale than with the JSSF scale. From Data C, correlations were significant between patient satisfaction and outcome for AH and HL by the JSSF scales and for AH, HL, and LT by the JOA scale. CONCLUSIONS: The validity of both scales was high. Clinical evaluation of the therapeutic results using these scales would be highly reliable.

Ankle↗

A quick and reliable screening measure for OCD in youth: reliability and validity of the obsessive compulsive scale of the Child Behavior Checklist.

BACKGROUND: The high prevalence and morbidity of obsessive compulsive disorder (OCD) in youth, the secretive nature of the disorder leading to under-recognition, and the lack of specialized child psychiatry services in many areas suggest that a simple, quick, and reliable screening tool to identify cases could be very useful to clinicians who work with children. METHOD: We used 8 items from the Child Behavior Checklist (CBCL), an empirically derived instrument free of clinician bias, to investigate the usefulness of a previously reported CBCL-based obsessive compulsive scale (OCS) by Nelson et al [Nelson EC, Hanna GL, Hudziak JJ, Botteron KN, Heath AC, Todd RD. Obsessive-compulsive scale of the Child Behavior Checklist: Specificity, sensitivity, and predictive power. Pediatrics 2001;108(1):E14] in a separate cohort of youth with OCD. We computed the psychometric properties of the OCS in our sample of youth with OCD and in psychiatric and normal controls, and compared these to the published values. RESULTS: Using the recommended cutoff between the 60th and 70th percentiles of the OCS to best predict the presence of OCD, we found very high sensitivity (92%-78%), specificity (86%-94%), negative predictive value (96%-90%), and positive predictive value (77%-86%). CONCLUSIONS: The OC scale of the CBCL shows good reliability and validity and acceptable psychometric properties to help discriminate youth with OCD.

Case-Control Studies↗

Determining reliable cognitive change after epilepsy surgery: development of reliable change indices and standardized regression-based change norms for the WMS-III and WAIS-III.

PURPOSE: Reliable change indices (RCIs) and standardized regression-based (SRB) change scores norms were established for the recently revised Wechsler Adult Intelligence Scale-III (WAIS-III) and Wechsler Memory Scale-III (WMS-III) in patients with complex partial seizures. Establishment of such standardized change scores can be useful in determining the effects of epilepsy surgery on cognitive functioning independent of test-retest artifacts including practice effects. METHODS: Forty-two nonoperated-on adult patients with complex partial seizures (primarily of temporal lobe onset) were administered the WMS-III and WAIS-III on two occasions (mean 7-month interval). All patients were receiving stable antiepileptic drug (AED) treatment at both testings. RCI and SRB change scores were calculated. Confidence interval cutoff scores (90% and 80%) and standardized regression equations were calculated for each of the WAIS-III and WMS-III Primary Indices and individual subtests. Age, gender, education, test-retest interval, preoperative test performance, seizure onset, and seizure duration were predictor variables for the SRB equations. RESULTS: Test-retest reliabilities for the WAIS-III and WMS-III Primary Indices were within acceptable ranges, although considerable individual subtest variability was found. Preoperative performance was the single largest contributor to each of the predictive regression equations. Age, gender, education, seizure onset, and seizure duration contributed modest variance to several of the regression equations. CONCLUSIONS: We calculated both RCI and SRB change score indices for the recently revised Wechsler instruments. These formulas help control for test-retest methodologic artifacts and provide a standardized method with which to examine both individual and group level cognitive change after epilepsy surgery.

Adult↗

Prodromal assessment with the structured interview for prodromal syndromes and the scale of prodromal symptoms: predictive validity, interrater reliability, and training to reliability.

As the number of studies related to the early identification of and intervention in the schizophrenia prodrome continues to grow, it becomes increasingly critical to develop methods to diagnose this new clinical entity with validity. Furthermore, given the low incidence of patients and the need for multisite collaboration, diagnostic and symptom severity reliability is also crucial. This article provides further data on these psychometric parameters for the prodromal assessment instruments developed by the Prevention through Risk Identification, Management, and Education (PRIME) prodromal research team at Yale University: the Structured Interview for Prodromal Syndromes and the Scale of Prodromal Symptoms. It also presents data suggesting that excellent interrater reliability can be established for diagnosis in a day-and-a-half-long training workshop.

Adolescent↗

Is the Beck Depression Inventory reliable over time? An evaluation of multiple test-retest reliability in a nonclinical college student sample.

The Beck Depression Inventory (BDI) is one of the most widely used measures of depression. Many studies have examined the reliability and validity of the BDI. However, we found no published studies that considered the stability of the BDI over multiple administrations (i.e., more than 3 trials), such as is common in clinical trials research and during some clinical interventions. The purpose of this study is to examine the multiple test-retest reliability of the BDI in a presumably nonclinical sample. Results show a 40% decline in BDI scores over 8 weeks, a main effect that accounts for approximately 10% of the variance. We achieved a 40% decrease in self-reported symptoms of depression due to repeated measurement alone, not due to any intervention. This change likely represents measurement error with this instrument rather than any "real" change in depression. The limitations of this study, its implications for research, and its applications to clinical practice are discussed.

Adult↗

[Reliability and accuracy of reported causes of death from cancer. I. Reliability of all cancer reported in the State of Rio de Janeiro, Brazil]

Mortality records are often used in epidemiological studies, particularly in cancer studies. This paper aims to evaluate reliability and accuracy of cancer mortality data in Rio de Janeiro, Brazil. A systematic random sample of 394 death certificates was obtained from a total of 12615 cancer deaths. This sample was recoded by an independent codifier. A kappa coefficient of 0.89 (95% C.I. 0.86-0.92) was obtained to the third digit, which increases to 0.95 (95% C.I. 0.94-0.96) when restricted to the mortality list used in international publications. The positive predictive value was 95.7% for this sample. These results reveal a high standard reliability of cancer mortality records in the State of Rio de Janeiro making them suitable for use in epidemiological research.

Journal Article↗

Reliability between nurse managers: the key to the high-reliability organization.

Flawless execution rests in the hands of nurse managers. No one can work alone in health care any more. We are interdependent and know that the best outcomes happen when practices are organized around collegial supportive structures rather than autonomous competitive units. We are only as strong as our weakest link. If all managers see the big picture and look beyond their units for what is right for the common good, we will achieve high-reliability organizations in health care. In turn health care organizations will become very safe places to operate. Shared governance structures for nurse managers are the perfect vehicle to develop collaborative organizations and flawless execution, and to adopt high-reliability organization principles.

Decision Making, Organizational↗

The reliability of reliability.

Forty-five original articles addressing the subject of examiner reliability were reviewed to determine if the findings were adequately substantiated by the statistical analyses and experimental designs employed by the authors. Only 10 studies were determined to have properly supported conclusions, while an additional three studies contained correct conclusions by coincidence. Eight investigations had invalid designs and three contained claims that were contradicted by the author's findings. Half the studies were found to have conclusions that were based on inappropriate or inconclusive statistical analysis. To date, the research presented in the chiropractic literature cannot substantiate claims concerning the reliability of any diagnostic instrumentation or palpatory procedures commonly employed by chiropractic physicians.

Chiropractic↗

Endodontic recall radiographs: how reliable is our interpretation of endodontic success or failure and what factors affect our reliability?

Three hundred thirty cases were selected from an endodontic practice. Postoperative and recall radiographs of each case were examined by four endodontists for an interpretation of treatment success or failure. One hundred eighteen cases were examined a second time by each endodontist. Initial analysis showed substantial inconsistency in both inter- and intraobserver interpretation. The cases were then categorized by average radiographic density differences within radiograph sets, anatomic location of the treated tooth, technical compatibility within radiograph sets, and by length of time between postoperative and recall radiographs. It appears that these factors do not affect reliability of success/failure interpretation.

Dental Pulp Cavity↗

The operons, a criterion to compare the reliability of transcriptome analysis tools: ICA is more reliable than ANOVA, PLS and PCA.

The number of statistical tools used to analyze transcriptome data is continuously increasing and no one, definitive method has so far emerged. There is a need for comparison and a number of different approaches has been taken to evaluate the effectiveness of the different statistical tools available for microarray analyses. In this paper, we describe a simple and efficient protocol to compare the reliability of different statistical tools available for microarray analyses. It exploits the fact that genes within an operon exhibit the same expression patterns. In order to compare the tools, the genes are ranked according to the most relevant criterion for each tool; for each tool we look at the number of different operons represented within the first twenty genes detected. We then look at the size of the interval within which we find the most significant genes belonging to each operon in question. This allows us to define and estimate the sensitivity and accuracy of each statistical tool. We have compared four statistical tools using Bacillus subtilis expression data: the analysis of variance (ANOVA), the principal component analysis (PCA), the independent component analysis (ICA) and the partial least square regression (PLS). Our results show ICA to be the most sensitive and accurate of the tools tested. In this article, we have used the protocol to compare statistical tools applied to the analysis of differential gene expression. However, it can also be applied without modification to compare the statistical tools developed for other types of transcriptome analyses, like the study of gene co-expression.

Bacillus subtilis↗

The standard error in the Jacobson and Truax Reliable Change Index: the classical approach to the assessment of reliable change.

Researchers and clinicians using Jacobson and Truax's index to assess the reliability of change in patients, or its counterpart by Chelune et al., which takes practice effects into account, are confused by the different ways of calculating the standard error encountered in the literature (see the discussion started in this journal by Hinton-Bayre). This article compares the characteristics of (1) the standard error used by Jacobson and Truax, (2) the standard error of difference scores used by Temkin et al. and (3) an adaptation of Jacobson and Truax's approach that accounts for difference between initial and final variance. It is theoretically demonstrated that the last variant is preferable, which is corroborated by real data.

Algorithms↗

How reliable are reliability studies of fracture classifications? A systematic review of their methodologies.

Two independent reviewers performed a search in MEDLINE and EMBASE for fracture classification reliability studies. Data were obtained on classifications, image modalities, fracture selection processes, sample sizes and their justification, type and number of raters, practical issues for the classification sessions, statistical methods, and results. A 10-item checklist was devised for quality assessment of methodologies. 44 studies assessing 32 fracture classification systems were included. We found a wide variation of methodologies. For instance, the median number of raters was 5 (2-36) and the median number of fractures was 50 (10-200). This selection was considered representative in 17/44 of the studies. The true distribution of classification categories was estimated in 9 studies. The kappa coefficient was mostly used (39/44) to quantify the raters' agreement. Methodological issues are discussed. Given limitations in the use and interpretation of kappa coefficients, investigators should consider alternative methods that focus upon the accuracy of the classification systems. The development and adoption of a systematic methodological approach to the development and validation of fracture classification systems is needed.

Fractures, Bone↗

Reliability of the AMDP-system. A preliminary report on a multicentre exercise on the reliability of psychopathological assessment.

The AMDP-System is a documentation system for psychiatric data widely in use in the German-speaking countries. A summary of results of a multicentered study of interrater agreement of the Psychopathology Scale is presented. A new index of rater agreement was tested and the notion is discussed that the judgement of the presence and the absence of a symptom are two different processes with different reliability.

Germany, West↗