PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data quality”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Improving data quality for environmental fate models: a least-squares adjustment procedure for harmonizing physicochemical properties of organic compounds.

Physicochemical properties (vapor pressure, aqueous solubility, octanol solubility, Henry's law constant, and octanol-air and octanol-water partition coefficients) and their temperature dependencies are required for fate modeling of environmental pollutants. To be internally consistent, measured values for these properties often must be adjusted. The goal of adjusting the property values for consistency is to more accurately estimate the true values. However, consistency and accuracy are not synonymous. If there are systematic errors in one property, then adjustment for consistency may reduce the accuracy of other property data. Here, we provide methods for achieving consistency and improving accuracy in the selection of partitioning properties from literature sources. First, we show that a widely used procedure does not always minimize the adjustments of property values derived from the literature when harmonizing them according to thermodynamic constraints. In such cases, the final adjusted values (FAVs) are unnecessarily different from the literature-derived values (LDVs) selected from measurements. We present an improved procedure based on the theory of least squares that minimizes the adjustment of LDVs and allows quantitative propagation of uncertainty from LDVs to FAVs. When this procedure is applied to partitioning properties for 30 organic chemicals, FAVs obtained differ by up to 30% from those calculated with the current adjustment procedure. Second, we point out that the adjustment procedure is only appropriate for correcting random errors in measurement data. Biased LDVs must be identified and corrected prior to harmonization. Using a set of 16 PCB congeners as a case study, we provide methods to identify biased data and discuss possible sources of bias. We present a new interpretation of property data for the PCBs and a new set of internally consistent properties and quantitative structure-property relationships that we recommend as the best currently available.

Environmental Pollutants↗

[Record linkage in the study of infant mortality: some aspects concerning data quality].

"The article concerns some aspects of the record linkage carried out by the ISTAT [Central Statistical Institute] for the study of infant mortality among the 1975 birth cohort in seven Italian regions." About eight percent of the records are found to be unlinked. A log-linear model is used to identify the variables affecting the percentage of unlinked records. (summary in ENG, FRE)

Data Collection↗

The Swedish SF-36 Health Survey--I. Evaluation of data quality, scaling assumptions, reliability and construct validity across general populations in Sweden.

We document the applicability of the SF-36 Health Survey, which was translated into Swedish using methods later adopted by the International Quality of Life Assessment (IQOLA) Project procedures. To test its appropriateness for use in Sweden, it was administered through mail-out/mail-back questionnaires in seven general population studies with an average response rate of 68%. The 8930 respondents varied by gender (48.2% men), age (range 15-93 years, mean age 42.7), marital status, education, socio-economic status, and geographical area. Psychometric methods used in the evaluation of the SF-36 in the U.S. were replicated. Over 90% of respondents had complete items for each of the eight SF-36 scales, although more missing data were observed for subjects 75 years and over. Scale scores could be computed for the vast majority of respondents (95% and over); slightly fewer in the oldest subgroup. Item-internal consistency was consistently high across socio-demographic subgroups and the eight scales. Most reliability estimates exceeded the 0.80 level. The highest reliability was observed for the Bodily Pain Scale where all subgroups met the 0.90 level recommended for individual comparisons; coefficients at or above 0.90 were also observed in most subgroups for the Physical Functioning Scale. Tests of scaling assumptions including hypothesized item groupings, which reflect the construct validity of scales, were consistently favorable across subgroups, although lower rates were noted in the oldest age group. In conclusion, these studies have yielded empirical evidence supporting the feasibility of a non-English language reproduction of the SF-36 Health Survey. The Swedish SF-36 is ready for further evaluation.

Adolescent↗

Quality of data in multiethnic health surveys.

OBJECTIVE: There has been insufficient research on the influence of ethno-cultural and language differences in public health surveys. Using data from three independent studies, the authors examine methods to assess data quality and to identify causes of problematic survey questions. METHODS: Qualitative and quantitative methods were used in this exploratory study, including secondary analyses of data from three baseline surveys (conducted in English, Spanish, Cantonese, Mandarin, and Vietnamese). Collection of additional data included interviews with investigators and interviewers; observations of item development; focus groups; think-aloud interviews; a test-retest assessment survey; and a pilot test of alternatively worded questions. RESULTS: The authors identify underlying causes for the 12 most problematic variables in three multiethnic surveys and describe them in terms of ethnic differences in reliability, validity, and cognitive processes (interpretation, memory retrieval, judgment formation, and response editing), and differences with regard to cultural appropriateness and translation problems. CONCLUSIONS: Multiple complex elements affect measurement in a multiethnic survey, many of which are neither readily observed nor understood through standard tests of data quality. Multiethnic survey questions are best evaluated using a variety of quantitative and qualitative methods that reveal different types and causes of problems.

Black or African American↗

Mitochondrial DNA control region sequences in Koreans: identification of useful variable sites and phylogenetic analysis for mtDNA data quality control.

We have established a high-quality mtDNA control region sequence database for Koreans. To identify polymorphic sites and to determine their frequencies and haplotype frequencies, the complete mtDNA control region was sequenced in 593 Koreans, and major length variants of poly-cytosine tracts in HV2 and HV3 were determined in length heteroplasmic individuals by PCR analysis using fluorescence-labeled primers. Sequence comparison showed that 494 haplotypes defined by 285 variable sites were found when the major poly-cytosine tract genotypes were considered in distinguishing haplotypes, whereas 441 haplotypes were found when the poly-cytosine tracts were ignored. Statistical parameters indicated that analysis of partial mtDNA control region which encompasses the extended regions of HV1 and HV2, CA dinucleotide repeats in HV3 and nucleotide position 16497, 16519, 456, 489 and 499 (HV1ex+HV2ex+HV3CA+5SNPs) and the analysis of another partial mtDNA control region including extended regions of HV1 and HV2, HV3 region and nucleotide position 16497 and 16519 (HV1ex+HV2ex+HV3+2SNPs) can be used as efficient alternatives for the analysis of the entire mtDNA control region in Koreans. Also, we collated the basic informative SNPs, suggested the important mutation motifs for the assignment of East Asian haplogroups, and classified 592 Korean mtDNAs (99.8%) into various East Asian haplogroups or sub-haplogroups. Haplogroup-directed database comparisons confirmed the absence of any major systematic errors in our data, e.g., a mix-up of site designations, base shifts or mistypings.

Asian People↗

The General Practice Assessment Survey (GPAS): tests of data quality and measurement properties.

OBJECTIVES: The aim of this study was to describe the psychometric properties of the General Practice Assessment Survey (GPAS) and its acceptability to patients in the UK. GPAS comprises seven multiple item scales and two single item scales addressing nine key areas of primary care activity (access, technical care, communication, inter-personal care, trust, knowledge of patient, nursing care, receptionists and continuity of care). A further four single items relate to patients' perceptions of the GP's role in referral and co-ordination of care, their willingness to recommend their GP and their overall satisfaction with care received. METHODS: Two hundred consecutive patients attending routine consulting sessions at 55 inner London practices were invited to complete the GPAS questionnaire. The acceptability, reliability and validity of GPAS was assessed using standard psychometric techniques. RESULTS: Out of 11 000 patients, 7247 (66%) completed a questionnaire in a GP surgery. Fifty-five out of a separate sample of 77 patients attending one practice completed a second questionnaire mailed to them 1 week following their attendance. GPAS was acceptable to patients as evidenced by low proportions of missing data for all items, and a full range of possible scores for all but one of the nine scales. Reliability of the instrument was good. Multiple item scales had excellent internal consistency, high item-total correlations, and test-retest reliability. Scaling assumptions were confirmed, with six of the seven scales achieving 100% scaling success (convergent and discriminant validity). Construct validity was evident, although this requires further evaluation against external measures. CONCLUSIONS: GPAS is a useful instrument for assessing several important dimensions of primary care. It is acceptable, reliable and valid, and has the potential for versatility in mode of administration. It will be a useful instrument for practices, primary care groups and primary care researchers evaluating key areas of primary care activity. Further work is required to evaluate its performance in non-inner-city settings and to evaluate further its validity against external criteria.

Adolescent↗

Data quality and age: health and psychobehavioral correlates of item nonresponse and inconsistent responses.

This study examined item nonresponse and inconsistent responses (IRs) and their health and psychobehavioral correlates in a population-based survey of adults 65 years and older. We administered an in-person questionnaire concerning physical, social, and psychological health to 1,155 men (mean age = 73.7 years) and 1,942 women (mean age = 74.8 years). Nonresponse rates varied with item topic, and "don't know" (DK) responses were more common than refusals. DKs increased with age of respondent, tended to be more common in women than men, and were associated with poorer physical, cognitive, and psychological functioning. Conversely, IRs increased with age among men but not women, but were also associated with poorer physical, cognitive, and psychological functioning. Results are discussed in terms of motivational and attentional factors, and their implications for survey research with the frail elderly and very old are noted.

Affect↗

A comparison of costs and data quality of three health survey methods: mail, telephone and personal home interview.

Three survey modes--a self-administered mailed questionnaire, a telephone interview, and a home interview--were assessed for survey costs, adequacy of completion, test-retest reliability, validity of responses to medical questions and estimates of morbidity. Costs per household for each mode were $A42.75, $A74.33, and $71.89, respectively. Item omission was confined virtually to the mail mode and averaged 5.5% over 84 questions assessed, while telephone and home interview modes averaged 0.4% and 0.2%, respectively. "Don't knows" were virtually absent for all questions except those about precise details (names, places, etc.) of events occurring often 10-15 years before the survey; no mode differences were observed. The mail mode produced less reliable responses to questions about environmental exposure to hazardous chemicals or activities when considered question-by-question, but differences were not significant among modes when all questions were grouped. Reliability was high to medical questions and no mode differences were observed. Medical conditions which would require a medical diagnosis for subjects to be able to report them were more reliably answered than conditions described in broad or lay terms. Validity of answers to medical questions varied across modes and types of questions; underreporting of medical conditions was highest in the mail mode and was lowest for conditions requiring a diagnosis. Overreporting was lowest in the mail mode and highest for conditions requiring a diagnostic opinion.

2,4,5-Trichlorophenoxyacetic Acid↗

Does nutritionist review of a self-administered food frequency questionnaire improve data quality?

OBJECTIVE: This study sought to evaluate the benefit of utilizing a nutritionist review of a self-administered food frequency questionnaire (FFQ), to determine whether accuracy could be improved beyond that produced by the self-administered questionnaire alone. DESIGN: Participants randomized into a dietary intervention trial completed both a FFQ and a 4-day food record (FR) at baseline before entry into the intervention. The FFQ was self-administered, photocopied and then reviewed by a nutritionist who used additional probes to help complete the questionnaire. Both the versions before nutritionist review and after nutritionist review - were individually compared on specific nutrients to the FR by means, correlations and per cent agreement into quintiles. SETTINGS AND SUBJECTS: Three hundred and twenty-four people, a subset of participants from the Polyp Prevention Trial - a randomized controlled trial examining the effect of a low-fat, high-fibre, high fruit and vegetable dietary pattern on the recurrence of adenomatous polyps - were recruited from clinical centres at the University of Utah, University of Buffalo, Memorial Sloan Kettering Cancer Center in New York and Kaiser Permanente Medical Program in Oakland. RESULTS: Reviewing the FFQ increased correlations with the FR for every nutrient, and per cent agreement into quintiles for all nutrients except calcium. Energy was underestimated in both versions of the FFQ but to a lesser degree in the version with review. CONCLUSIONS: One must further evaluate whether the increases seen with nutritionist review of the FFQ will enhance our ability to predict diet-disease relationships and whether it is cost-effective when participant burden and money spent utilizing trained personnel are considered.

Data Collection↗

Monitoring the integration of hospital information systems: How it may ensure and improve the quality of data.

Integration of hospital departmental information systems (HDIS) has become a common but difficult issue. In May 2003, the Department of Biostatistics and Medical Informatics implemented a Virtual Electronic Patient Record (VEPR) for the Hospital S. João (HSJ), a university hospital with over 1350 beds. The system integrates clinical data from 10 legacy HDIS plus the Hospital Administrative Database (HAD), aiming to deliver all patient information to health professionals. Currently, around 500 medical doctors use the system on a regular basis and the HSJ-VEPR retrieves an average of 3,000 new reports per day, in PDF or HTML formats. This paper describes and discusses the role of monitoring in the assurance and improvement of data quality. Three approaches were put in place: (a) monitoring the HSJ-VEPR concerning the frequency of clinical records retrieved from the DIS by checking if the daily number of reports sent by the HDIS fell in the normal range from similar week days; (b) monitoring inconsistencies in the patient's identification by cross-checking between HDIS and HAD; and (c) monitoring the integrity of clinical records delivered to medical doctors through the HSJ-VEPR by checking their digital signature. During 2005, the monitoring system detected 53 unusual frequency patterns of which 44 corresponded to real problems. Over a 6 months period, more than 400 alerts were generated concerning inconsistencies in the patient's identification found in laboratory reports. Nevertheless, a significant reduction in the number of these inconsistencies occurred - from 116 in July to 10 in December 2005--due to implementation of preventive measures by the DIS. Finally, report's integrity was checked each time the report was asked to be visualized i.e. in more than one hundred thousand times during a one year period. In conclusion, all information available in hospital information systems can and should be used to trigger alerts of malfunctions and inconsistencies, in order to improve data quality and ensure a better health care.

Hospital Information Systems↗

Translation and performance of the Norwegian SF-36 Health Survey in patients with rheumatoid arthritis. I. Data quality, scaling assumptions, reliability, and construct validity.

The SF-36 was translated into Norwegian following the procedures developed by the International Quality of Life Assessment (IQOLA) Project. To test for the appropriateness of the Norwegian Version 1.1 of the SF-36 in patients with rheumatoid arthritis (RA), 1552 RA patients were mailed the form. Psychometric methods used in previous U.S. and Swedish studies were replicated. The response rate was 66%. The sample (mean age 62 years, mean disease duration 13 years) was over-represented by females (79%). Totally, 74% of the questionnaires were complete. Missing value rates per item ranged from 0.4% to 9.0% (mean 4.2%). In the Role-Emotional scale, all three items had missing value rates above average and higher than reported in the U.S. and Swedish studies. Tests of scaling assumptions confirmed the hypothesized structure of the questionnaire, but results were suboptimal in the General Health scale. In all scales the Cronbach's alphas exceeded the 0.70 standard for group comparisons. In the Physical Functioning scale, Cronbach's alpha exceeded the 0.90 standard for individual comparisons. There was good evidence for the construct validity of the questionnaire. Generally, the Norwegian SF-36 version 1.1 distributed to RA patients held the psychometric properties found in other countries and in normal populations. The translations of items in the General Health and Role-Emotional scales were reassessed. Minor deficiencies were detected and changed (SF-36 Norwegian Version 1.2).

Arthritis, Rheumatoid↗

2D mapping by Kohonen networks of the air quality data from a large city.

The 15-variable environmental data (7 concentrations: CO, SO2, O3, NOx, NO, NO2, particulate matter smaller than 10 micron (PM10), and 8 weather data: cloudiness, rainfall, insolation factor (Isfi), temperature, pressure at two locations, and wind intensity with direction) in a period of 45 days with 1-h intervals were extracted from a larger database of concentrations recorded in minute intervals for the same time period. The monitoring site was located in the City of Buenos Aires in a relatively heavy traffic crossroad of two avenues. The data required special pretreatment where the hourly content of rain, wind intensity, wind velocity, and cloudiness were concerned. The new variable named insolation factor (relative UV radiation) calculated on the basis of the general meteorological data, the geographic position of the monitoring site, cloudiness, date, and the time of the recording was composed. The relative intensity of UV radiation was modeled by a Gaussian function, multiplied by a cloudiness factor. Based on the 14-variable input and the 1-variable output (ozone) data, first, the clustering of all 980 data records was made. The top map clustering showing the ozone concentration was related to the maps of all 14 variables. The link between O3 clusters, NO2, and Isfi weight levels is shown and discussed. As a preliminary result of this study some of the most interesting correlations between the maps and remaining variables are given.

Journal Article↗