PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Individual reliability of amplitude distribution in topographical mapping of EEG.

Whilst there is an accumulation of evidence suggesting that many quantitative EEG parameters show good stability and reliability, no previous study has considered whether the spatial distribution of EEG amplitude is reliable over time within a session. This study reports on the spatio-temporal reliability of EEG using data recorded from 24 subjects in a baseline condition with eyes open and also whilst performing a simple motor task. Both the internal stability and test-retest reliability for electrode parameters were comparable to previously published data. For most individuals, amplitude distribution was stable within each recording condition, but the test-retest reliability after 40 min was less good with the poorest reliability in the delta frequency band. Most subjects showed spatio-temporal reliability of less than 0.7 in at least one frequency band. In contrast, spatio-temporal reliability for the group average was good and exceeded 0.88 in all frequency bands. It is argued that the results indicate that reliability is insufficient to allow topographical comparisons for a single individual, but is more than adequate to allow group comparisons.

Adolescent↗

Test-retest reliability of adult surveillance measures for physical activity and inactivity.

BACKGROUND: Several physical activity measures used in U.S. surveillance systems lack estimates of reliability in this country. This information is needed among diverse populations of women and men, to aid in interpretation and use of the measures. The objective of this study was to document the test-retest reliability of several measures of physical activity and inactivity used in surveillance in a diverse adult population. METHODS: Test and retest surveys were conducted over the telephone with 106 African-American and white women and men living in Forsyth County, North Carolina or Jackson, Mississippi in 2003. Physical activity and inactivity were self-reported using surveillance measures, such as from the Behavioral Risk Factor Surveillance System. Reliability was determined using kappa and intraclass correlation coefficients (ICCs) overall and separately by gender and race. RESULTS: Thirteen percent of the participants met recommendations for physical activity, 44% were insufficiently active, and 43% were inactive. Reliability of the measures to categorize participants into these categories was 0.44 (95% confidence interval [CI]=0.27-0.58). The reliability of several categoric definitions of leisure activity ranged from 0.46 to 0.68. Occupational activity had substantial reliability (0.82, 95% CI=0.72-0.89), while any transportation activity (0.27, 95% CI=0.09-0.44) and walking (0.40, 95% CI=0.23-0.55) were lower. Indicators of inactivity categorized at >7 hours/week included time per week on the computer (0.83, 95% CI=0.57-0.78) and time per week watching television (0.40, 95% CI=0.22-0.54). Some gender and racial differences were noted in the reliability estimates. CONCLUSIONS: In conclusion, this study provides estimates of test-retest reliability for several physical activity and inactivity measures used for surveillance purposes. Validity data, coupled with the reliability estimates reported here and elsewhere, can aid in interpretation and use of these measures in surveillance, as well as in epidemiologic studies.

Black or African American↗

Intra- and interrater reliability of the Ergo-Kit functional capacity evaluation method in adults without musculoskeletal complaints.

OBJECTIVE: To evaluate the intra- and interrater reliability of tests from the Ergo-Kit (EK) functional capacity evaluation method in adults without musculoskeletal complaints. DESIGN: Within-subjects design. SETTING: Academic medical center in The Netherlands. PARTICIPANTS: Twenty-seven subjects without musculoskeletal complaints (15 men, 12 women). INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: Seven EK tests (2 isometric, 3 dynamic lifting, 2 manipulation tests) were each assessed 3 times (over 4 days), twice by 1 rater (R1) and once by another rater (R2). Intrarater reliability was calculated using the EK test scores assessed by R1. Interrater reliability was calculated using the EK test scores assessed by both raters. Counterbalancing the rater order made possible the calculation of 2 interrater reliability levels (at time intervals of 4 and 8d). All reliability levels were expressed as intraclass correlation coefficients (ICCs). RESULTS: Intrarater and interrater reliability (8-d time interval) was high (ICC, >.80) for the isometric lifting tests, moderate (ICC range, .50-.80) for the dynamic lifting tests, and low (ICC, <.50) for the manipulation tests. The interrater reliability of the isometric and dynamic lifting tests (4-d time interval) was high (ICC, >.80), and it was moderate (ICC range, .50-.80) for both manipulation tests. CONCLUSIONS: The isometric and dynamic lifting tests of the EK have a moderate to high level of reliability; the manipulation tests have a low level of reliability.

Adult↗

Test-retest and inter-rater reliability for the Comprehensive Developmental Inventory for Infants and Toddlers diagnostic and screening tests.

BACKGROUND: Reliability information for the Comprehensive Developmental Inventory for Infants and Toddlers diagnostic (CDIITDT) and screening tests (CDIITST) is inadequate. AIM: To assess the test-retest and inter-rater reliability of the CDIITDT and CDIITST. STUDY DESIGN: A repeated measures design was selected. SUBJECTS: Non-disabled term (n=15; mean age 8.4+/-1.6 months) and preterm infants (n=16; mean age 9.3+/-2.9 months), and children with developmental disabilities (n=15; mean age 24.7+/-11.8 months) were recruited. A single rater assessed the children twice in 3 days to examine the test-retest reliability; and a second rater observed and scored performance while the same rater conducted the first assessment for the inter-rater reliability analysis. OUTCOME MEASURES: The raw score, developmental age (DA) and developmental quotient (DQ)/Z score for the six subtests, two motor subdomains and the whole test were used as outcome measures for the CDIITDT and CDIITST. RESULTS: The test-retest reliabilities for the CDIITDT were rated good for the three pediatric groups (ICC 0.76-1.00), with the exception of moderate ratings for the self-help subtest for the term infants and for the social, self-help and fine-motor DQs for the preterm group. The CDIITDT inter-rater reliabilities were good for the three groups (ICC 0.76-1.00), with the exception of only moderate reliability for the cognitive DQs for the preterm infants. The reliabilities for the whole CDIITST for the three groups were high (ICC 0.93-1.00). CONCLUSION: The reliabilities for the whole CDIITDT and its various subtests and the whole CDIITST are acceptable for clinical use.

Child↗

Reliability of the "Sydney," "Sunnybrook," and "House Brackmann" facial grading systems to assess voluntary movement and synkinesis after facial nerve paralysis.

OBJECTIVE: To investigate the extent of within-system reliability and between-system correlation for the "Sydney" and "Sunnybrook" systems of grading facial nerve paralysis, and to examine the interobserver reliability and agreement of the "House Brackmann" grading system. STUDY DESIGN: A fixed-effects reliability study in which 6 otolaryngologists viewed videotapes of patients with facial nerve paralysis. SETTING: University and medical Centers. PATIENTS: Patients with unilateral lower motor neurone facial nerve dysfunction greater than 1 year after onset, none of whom had undergone surgical reanimation procedures. Intervention Twenty-one patients with facial nerve paralysis were videotaped while they performed a protocol of facial movements. Six otolaryngologists viewed the videotapes and scored them with the Sydney and Sunnybrook systems, and then gave a House Brackmann grade. MAIN OUTCOME MEASURE: The 3 systems of grading facial nerve paralysis were evaluated and compared with the use of intraclass correlation coefficients, Pearson's weighted kappa, and percentage exact agreement values. RESULTS: The Sydney and the Sunnybrook systems had good intrasystem reliability and high intersystem association for the assessment of voluntary movement. Grading of synkinesis was found to have low reliability both within and between systems. The House Brackmann system had substantial reliability as shown by weighted kappa but had a percentage exact agreement of 44%. CONCLUSIONS: For clinical grading of voluntary movement, there is good correlation between ratings given on the Sydney and Sunnybrook systems, and within each system there is good reliability. The assessment of synkinesis was far less reliable within, and less related between, systems. Although the reliability of the House Brackmann system was found to be high, examination of individual grades revealed some wide variation between trained observers.

Adult↗

Reliability of clinical temporomandibular disorder diagnoses.

Temporomandibular disorders (TMD) diagnoses can be viewed as the most useful clinical summary for classifying subtypes of TMD. The Research Diagnostic Criteria for TMD (RDC/TMD) is the most widely used TMD diagnostic system for conducting clinical research. It has been translated into 18 languages and is used by a consortium of 45 RDC/TMD-based international researchers. While reliability of RDC/TMD signs and symptoms of TMD has been amply reported, the reliability of RDC/TMD diagnoses has not. The aim of the study was to determine the reliability of clinical TMD diagnoses using standardized methods and operational definitions contained in the Research Diagnostic Criteria for Temporomandibular Disorders (RDC/TMD). Data came from reliability assessment trials conducted at 10 international clinical centers, involving 30 clinical examiners assessing 230 subjects. Intraclass correlation coefficients (ICC) were calculated to characterize the reliability. The reliability of the diagnoses was fair to good. Median ICCs for the diagnoses myofascial pain with and without limited opening were 0.51 and 0.60, respectively. Median ICC for arthralgia was 0.47 and 0.61 for disc displacement with reduction. RDC/TMD diagnoses of disc displacement without reduction, osteoarthritis and osteoarthrosis were not prevalent enough to calculate ICC's, but percent agreement was always >95%. The reliability of diagnostic classification improved when diagnoses were grouped into pain versus non-pain diagnoses (ICC=0.72) and for detecting any diagnosis versus no diagnosis (ICC=0.78). In clinical decision-making and research, arriving at a reliable diagnosis is critical in establishing a clinical condition and a rational approach to treatment. The RDC/TMD demonstrates sufficiently high reliability for the most common TMD diagnoses, supporting its use in clinical research and decision making.

Algorithms↗

Assuring the reliability of resident performance appraisals: more items or more observations?

BACKGROUND: The tendency to add items to resident performance rating forms has accelerated due to new ACGME competency requirements. This study addresses the relative merits of adding items versus increasing number of observations. The specific questions addressed are (1) what is the reliability of single items used to assess resident performance, (2) what effect does adding items have on reliability, and (3) how many observations are required to obtain reliable resident performance ratings. METHODS: Surgeon ratings of resident performance were collected for 3 years. The rating instrument had 3 single items representing clinical performance, professional behavior, and comparisons to other house staff. Reliability analyses were performed separately for each year, and variance components were pooled across years to compute overall reliability coefficients. RESULTS: Single-item resident performance rating scales were equivalent to multiple-item scales using conventional reliability standards. Increasing the number of rating items had little effect on reliability. Increasing the number of observations had a much larger effect. CONCLUSIONS: Program directors should focus on increasing the number of observations per resident to improve performance sampling and reliability of assessment. Increasing the number of rating items had little effect on reliability and is unlikely to assess new ACGME competencies adequately.

Clinical Competence↗

Reliability of clinicians versus radiologists for detecting abnormalities on hysterosalpingogram films.

OBJECTIVE: To evaluate the consistency of the identification of abnormal findings on hysterosalpingogram (HSG) and compare the reliability of clinicians to that of radiologists. DESIGN: Evaluation of reliability of diagnostic test. PATIENT(S): Women undergoing evaluation for infertility.INTEVENTION(S): Retrospective review of 50 HSG films by three reproductive endocrinologists and three radiologists. Each film was reread 30 days later in a blinded fashion. MAIN OUTCOME MEASURE(S): The consistency of each individual reader, the reliability of detecting specific abnormalities, and the consistency of clinicians compared with radiologists was evaluated with a kappa (K) statistic and interclass correlation coefficient (ICC). RESULT(S): Average intrareader reliability was high for the detection of normal uterus, normal tubes, and tubal obstruction and low for the detection of hydrosalpinx, uterine adhesions, and pelvic adhesions. Inter-reader reliability was high in the detection of normal uterine contour, normal tubal patency, and uterine filling defect and lower for the detection of a hydrosalpinx. The reliability of detecting pelvic adhesion or salpingitis isthmica nodosa was poor. CONCLUSION(S): Intrareader reliability was generally good, especially for the detection of normal findings. Agreement among different readers is lower in detecting rare outcomes such as hydrosalpinx and pelvic adhesion and salpingitis isthmica nodosa. Clinicians more reliably diagnose hydrosalpinx and tubal obstruction, while radiologists more reliably detect the more subtle findings of salpingitis isthmica nodosa or uterine adhesions.

Fallopian Tube Diseases↗

Predicting reliable regions in protein alignments from sequence profiles.

For applications such as comparative modelling one major issue is the reliability of sequence alignments. Reliable regions in alignments can be predicted using sub-optimal alignments of the same pair of sequences. Here we show that reliable regions in alignments can also be predicted from multiple sequence profile information alone. Alignments were created for a set of remotely related pairs of proteins using five different test methods. Structural alignments were used to assess the quality of the alignments and the aligned positions were scored using information from the observed frequencies of amino acid residues in sequence profiles pre-generated for each template structure. High-scoring regions of these profile-derived alignment scores were a good predictor of reliably aligned regions. These profile-derived alignment scores are easy to obtain and are applicable to any alignment method. They can be used to detect those regions of alignments that are reliably aligned and to help predict the quality of an alignment. For those residues within secondary structure elements, the regions predicted as reliably aligned agreed with the structural alignments for between 92% and 97.4% of the residues. In loop regions just under 92% of the residues predicted to be reliable agreed with the structural alignments. The percentage of residues predicted as reliable ranged from 32.1% for helix residues to 52.8% for strand residues. This information could also be used to help predict conserved binding sites from sequence alignments. Residues in the template that were identified as binding sites, that aligned to an identical amino acid residue and where the sequence alignment agreed with the structural alignment were in highly conserved, high scoring regions over 80% of the time. This suggests that many binding sites that are present in both target and template sequences are in sequence-conserved regions and that there is the possibility of translating reliability to binding site prediction.

Algorithms↗

Preliminary study of the reliability of assessment procedures for indications for chiropractic adjustments of the lumbar spine.

OBJECTIVE: To assess the intraexaminer and interexaminer reliability of clinicians trained in flexion-distraction technique to determine the need for chiropractic adjustment of each segment of the lumbar spine. DESIGN: This was an intraexaminer and interexaminer reliability study of commonly used chiropractic assessment procedures, including static and motion palpation and visual observation. SETTING: Chiropractic college; by four licensed chiropractors trained in flexion-distraction technique, two with more than 20 years' experience and two with 3 or fewer years' experience. SUBJECTS: Subjects were 18 volunteers; 16 were symptom free, and 2 had low back pain at the time the study was conducted. MAIN OUTCOME MEASURE: The kappa statistic was computed for all comparisons and interpreted in categories ranging from "poor" (<0.00) to "almost perfect" (>0.80). RESULTS: Intraexaminer reliability was greater than interexaminer reliability. For intraexaminer reliability there was considerable variation by segment and among the four examiners, but intraexaminer reliability appeared generally higher than interexaminer reliability. Overall, more subluxations were identified on the second examination than on the first. For interexaminer reliability, kappa scores were generally in the "poor" to "slight" categories. DISCUSSION: The results of this study, similar to those of other studies, indicate that even chiropractors trained in the same technique seem to show little consensus on the indications for the necessity to adjust specific segments of the spine. A more standardized assessment approach might be helpful in improving the reliability of diagnostic assessments.

Adult↗

Reliability of a comorbidity measure: the Index of Co-Existent Disease (ICED).

The reliability of an established comorbidity index (the Index of Co-Existent Disease) was tested using retrospective data from the case notes of elderly patients who had undergone total hip replacement. Inter-rater reliability was examined twice, first with two raters (n = 39) and then with three (n = 49). Intra-rater reliability was assessed using one rater (n = 45). Reasons for any lack of reliability were explored. The inter-rater reliability of the ICED was moderate (kappa 0.5-0.6). While the Functional Severity index performed well (kappa 0.6-1.0), the Index of Disease Severity subindex was less reliable (kappa 0.4-0.5). Differences between raters had an impact on the observed association between comorbidity and serious post-operative complications. Intra-rater reliability was excellent (kappa 0.9). Several reasons why inter-rater reliability was only moderate were identified, mostly related to uncertainties in applying the ICED. The reliability of the ICED needs to be improved before it is used more widely with retrospective data. This might be achieved by further clarification of the instructions for its use.

Activities of Daily Living↗

Reliability of radiographic assessment of acromial morphology.

The most widely used radiographic classification system for acromial morphology identifies three distinct acromial shapes: type I (flat), type II (curved), and type III (hooked). The purpose of this study was to measure the interobserver and intraobserver reliability of determinations of acromial morphology as defined by this system. Between 1990 and 1992, one hundred twenty-six supraspinatus outlet radiographs were obtained from 126 patients by technicians from Triangle Orthopaedic Associates in Durham, N.C. Six fellowship-trained shoulder surgeons independently reviewed each radiograph and classified it as type I, II, or III on the basis of established guidelines. Two surgeons classified each film a second time in random order. Analysis of variance was performed to obtain coefficients for interobserver and intraobserver reliability. Consensus ratings were then used to classify the 126 radiographs into consensus type I, consensus type II, or consensus type III groups. Percentages of type I, II, and III individual ratings within each consensus group were determined. The intraobserver reliability coefficient was 0.888, interpreted as good to excellent reliability. The interobserver reliability coefficient was 0.516, interpreted as poor to fair reliability. Of the 126 radiographs, 26 (20.6%) were rated as consensus type I, 76 (60.3%) were rated as consensus type II, and 24 (19.1%) were rated as consensus type III. The reliability of observer ratings was lowest when delineation between acromial types II and III was required. The low interobserver reliability makes comparisons of studies by different authors difficult to interpret and obscures the true incidence of acromial morphologic types. It also questions reported correlations between acromial type and shoulder pathologic conditions. It is concluded that a system that incorporates more objective classification criteria and acknowledges the continuous nature of acromial morphologic types may improve interobserver reliability and validate the system's use in making clinical and surgical judgments.

Acromion↗

Reliability of a quantification imaging system using magnetic resonance images to measure cartilage thickness and volume in human normal and osteoarthritic knees.

OBJECTIVE: The aim of this study was to evaluate the reliability of a software tool that assesses knee cartilage volumes using magnetic resonance (MR) images. The objectives were to assess measurement reliability by: (1) determining the differences between readings of the same image made by the same reader 2 weeks apart (test-retest reliability), (2) determining the differences between the readings of the same image made by different readers (between-reader agreement), and (3) determining the differences between the cartilage volume readings obtained from two MR images of the same knee image acquired a few hours apart (patient positioning reliability). METHODS: Forty-eight MR examinations of the knee from normal subjects, patients with different stages of symptomatic knee osteoarthritis (OA), and a subset of duplicate images were independently and blindly quantified by three readers using the imaging system. The following cartilage areas were analyzed to compute volumes: global cartilage, medial and lateral compartments, and medial and lateral femoral condyles. RESULTS: Between-reader agreement of measurements was excellent, as shown by intra-class correlation (ICC) coefficients ranging from 0.958 to 0.997 for global cartilage (P<0.0001), 0.974 to 0.998 for the compartments (P<0.0001), and 0.943 to 0.999 for the condyles(P<0.0001). Test-retest reliability of within-reader data was also excellent, with Pearson correlation coefficients ranging from 0.978 to 0.999 (P<0.0001). Patient positioning reliability was also excellent, with Pearson correlation coefficients ranging from 0.978 to 0.999 (P<0.0001). CONCLUSIONS: The results of this study establish the reliability of this MR imaging system. Test-retest reliability, between-reader agreement, and patient positioning reliability were all extremely high. This study represents a first step in the overall validation of an imaging system designed to follow progression of human knee OA.

Adult↗

Reliability and responsiveness of the Barry-Albright Dystonia Scale.

The reliability and responsiveness of the Barry-Albright Dystonia (BAD) Scale, a 5-point ordinal severity scale for secondary dystonia, was assessed. For interrater reliability, 13 raters scored 10 videotaped patients; for intrarater reliability, two raters rated the videotape again. For test-retest reliability, patients were rated on two occasions. Four inexperienced raters scored patients, received training, then scored additional patients. To assess responsiveness, we compared patient and physician global ratings of change (better, same, and worse) with BAD Scale score changes for 18 patients on intrathecal baclofen (ITB) trials. We assessed reliability with the intraclass correlation coefficient (ICC). The mean ICC for total BAD Scale scores were as follows: interrater reliability 0.866, intrarater reliability 0.967 and 0.978, test-retest reliability 0.978 (before training) and 0.967 (after training). We found the BAD Scale responsive to change, with most improved scores in patients rated by the patient, family, and neurosurgeon as 'better'. The total scores were reliable for experienced raters. We recommend training for clinicians interested in using the scale.

Adolescent↗

Reliability of a lifetime history of major depression: implications for heritability and co-morbidity.

BACKGROUND: In unselected samples, the diagnosis of major depression (MD) is not highly reliable. It is not known if occasion-specific influences on reliability index familial risk factors for MD, or how reliability is associated with risk for co-morbid anxiety disorders. METHODS: An unselected sample of 847 female twin pairs was interviewed twice, 5 years apart, about their lifetime history (LTH) of MD, generalized anxiety disorder (GAD) and panic disorder (PD). Familial influences on reliability were examined using structural equation models. Logistic regression was used to identify clinical features that predict reliable diagnosis. Co-morbidity was characterized using the continuation ratio test. RESULTS: The reliability of a LTH of MD over 5 years was fair (kappa = 0.43). There was no evidence for occasion-specific familial influences on reliability, and heritability of reliably diagnosed MD was estimated at 66%. Subjects with unreliably diagnosed MD reported fewer symptoms and, if diagnosed with MD only at the first interview, less impairment and help seeking, or, if diagnosed with MD only at the second interview, fewer episodes and a longer illness. A history of co-morbid GAD or PD is more prevalent among subjects with reliably diagnosed MD. CONCLUSIONS: A diagnosis of MD based on a single psychiatric interview incorporates a substantial amount of measurement error but there is no evidence that transient influences on recall and diagnosis index familial risk for MD. Quantitative indices of risk for MD based on multiple interviews should reflect both the characteristics of MD and the temporal order of positive diagnoses.

Anxiety Disorders↗

The wheelchair circuit: reliability of a test to assess mobility in persons with spinal cord injuries.

OBJECTIVE: To assess the reliability of a 9-task wheelchair circuit. DESIGN: Three test trials per subject were conducted by 2 raters. Inter- and intrarater reliability were examined. SETTING: Eight rehabilitation centers in the Netherlands. PARTICIPANTS: Convenience sample of 27 patients (age, >or=18 y) with spinal cord injury (SCI), all of whom were in the final stage of their inpatient rehabilitation. INTERVENTION: A wheelchair circuit was developed to assess mobility in subjects with SCI. The circuit consisted of 9 tasks: figure-of-8 shape, doorstep crossing, mounting a platform, sprint, walking, driving up treadmill slopes of 3% and 6%, wheelchair driving and transfer. MAIN OUTCOME MEASURE: Task feasibility, task performance time, and peak heart rates. RESULTS: The number of tasks that subjects could perform varied from 3 to 9. Feasibility intrarater reliability was.98, and the interrater reliability intraclass correlation coefficient (ICC) was.97. Performance time ICCs ranged from.70 to.99 (mean,.88) for intrarater reliability and from.76 to.98 (mean,.92) for interrater reliability. Heart rate ICCs ranged from.64 to.96 (mean,.81) for intrarater reliability and from.82 to.99 (mean,.89) for interrater reliability. CONCLUSIONS: The reliability of the wheelchair circuit was good. More research is needed to assess test validity and responsiveness.

Adult↗

Attainment and maintenance of reliability of axis I and II disorders over the course of a longitudinal study.

The baseline interrater reliability, test-retest reliability, follow-up interrater reliability, and follow-up longitudinal reliability of axis I and axis II diagnoses were assessed using the Structured Clinical Interview for DSM-III-R Axis I Disorders (SCID-I) and the Diagnostic Interview for DSM-III-R Personality Disorders (DIPD-R). Excellent kappas (>.75) were found in each of these reliability substudies for the majority of axis II disorders diagnosed five times or more. Dimensional reliability figures for axis II diagnoses were generally somewhat higher than those for their categorical counterparts; most intraclass correlation coefficients (ICCs) were in the excellent range. Excellent kappas were also found in each of these four reliability substudies for over half of the axis I disorders diagnosed five times or more. Taken together, the results of this study suggest that the reliability of axis II disorders is both good to excellent and practically equivalent to that found for most axis I disorders. The results of this study also suggest that high levels of reliability, once achieved, can be maintained over time for both axis I and II disorders.

Adult↗

Reliability of a multidimensional questionnaire for adults with treated complete cleft lip and palate.

The purpose of this study was to evaluate the reliability of a multidimensional questionnaire for Swedish adults with treated complete unilateral or bilateral cleft lip and palate (CLP). The questionnaire was designed to be used in the evaluation of adults with treated CLP after treatment. Before any conclusions were drawn from the results of the study we assessed the test-retest reliability of the questionnaire. The questionnaire included 168 questions and assessed the following domains: aesthetics, functions associated with CLP, satisfaction with treatment and perceived need for treatment, quality of life, depression and non-specific physical symptoms, body image, and jaw function. The subjects answered the questionnaire twice at a 2-3-week interval. Sixty-one adults (38 men, 23 women) mean age 24 years (range 20-29) participated in the study. The response rate for the questionnaire was acceptable at 75%. The test-retest reliability varied among the different domains. The reliability of questions regarding aesthetics, functions associated with CLP, and treatment satisfaction was good to excellent (intraclass correlation coefficient (ICC) = 0.51 to 0.89). Good to excellent (ICC = 0.61 to 1.0) reliability was also found for the quality of life in various life domains and the wellbeing scales. The reliability of the body image scale was moderate (kappa = 0.43-0.60) for most items and lower than that of other scales used in this study. The reliability of the mean depression symptom score (ICC = 0.93) and the mean non-specific physical symptoms score (ICC = 0.85) were excellent. The reliability of the mandibular function impairment was good (ICC = 0.67). The conclusion of the study is that an overall reliability was good for the multidimensional questionnaire.

Adult↗