PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data quality”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

An assessment of the data quality for NHEXAS--Part I: Exposure to metals and volatile organic chemicals in Region 5.

A National Human Exposure Assessment Survey (NHEXAS) was performed in U.S. Environmental Protection Agency (U.S. EPA) Region V, providing population-based exposure distribution data for metals and volatile organic chemicals (VOCs) in personal, indoor, and outdoor air, drinking water, beverages, food, dust, soil, blood, and urine. One of the principal objectives of NHEXAS was the testing of protocols for acquiring multimedia exposure measurements and developing databases for use in exposure models and assessments. Analysis of the data quality is one element in assessing the performance of the collection and analysis protocols used in NHEXAS. In addition, investigators must have data quality information available to guide their analyses of the study data. At the beginning of the program quality assurance (QA) goals were established for precision, accuracy, and method quantification limits. The assessment of data quality was complicated. First, quality control (QC) data were not available for all analytes and media sampled, because some of the QC data, e.g., precision of duplicate sample analysis, could be derived only if the analyte was present in the media sampled in at least four pairs of sample duplicates. Furthermore, several laboratories were responsible for the analysis of the collected samples. Each laboratory provided QC data according to their protocols and standard operating procedures (SOPs). Detection limits were established for each analyte in each sample type. The calculation of the method detection limits (MDLs) was different for each analytical method. The analytical methods for metals had adequate sensitivity for arsenic, lead, and cadmium in most media but not for chromium. The QA goals for arsenic and lead were met for all media except arsenic in dust and lead in air. The analytical methods for VOCs in air, water, and blood were sufficiently sensitive and met the QA goals, with very few exceptions. Accuracy was assessed as recovery from field controls. The results were excellent (> or = 98%) for metals in drinking water and acceptable (> or = 75%) for all VOCs except o-xylene in air. The recovery of VOCs from drinking water was lower, with all analytes except toluene (98%) in the 60-85% recovery range. The recovery of VOCs from drinking water also decreased when comparing holding times of < 8 and > 8 days. Assessment of the precision of sample collection and analysis was based on the percent relative standard deviation (% RSD) between the results for duplicate samples. In general, the number of duplicate samples (i.e., sample pairs) with measurable data were too few to assess the precision for cadmium and chromium in the various media. For arsenic and lead, the precision was excellent for indoor, and outdoor air (< 10% RSD) and, although not meeting QA goals, it was acceptable for arsenic in urine and lead in blood, but showed much higher variability in dust. There were no data available for metals in water and food to assess the precision of collection and analysis.

Adolescent↗

Data quality assurance for thermophysical property databases--applications to the TRC SOURCE data system.

To a significant degree processes of database development are based upon human activities, which are susceptible to various errors. Propagation of errors in the processing leads to a decrease in the value of original data as well as that of any database products. Data quality is a critical issue that every database producer must handle as an inseparable part of the database management. Within the Thermodynamics Research Center (TRC), a systematic approach to implement database integrity rules was established through the use of modern database technology, statistical methods, and thermodynamic principles. The four major functions of the system--error prevention, database integrity enforcement, scientific data integrity protection, and database traceability--are detailed in this paper.

Journal Article↗

Statistical issues in reporting quality data: small samples and casemix variation.

PURPOSE: To present two key statistical issues that arise in analysis and reporting of quality data. SUMMARY: Casemix variation is relevant to quality reporting when the units being measured have differing distributions of patient characteristics that also affect the quality outcome. When this is the case, adjustment using stratification or regression may be appropriate. Such adjustments may be controversial when the patient characteristic does not have an obvious relationship to the outcome. Stratified reporting poses problems for sample size and reporting format, but may be useful when casemix effects vary across units. Although there are no absolute standards of reliability, high reliabilities (interunit F > or = 10 or reliability > or = 0.9) are desirable for distinguishing above- and below-average units. When small or unequal sample sizes complicate reporting, precision may be improved using indirect estimation techniques that incorporate auxiliary information, and 'shrinkage' estimation can help to summarize the strength of evidence about units with small samples. CONCLUSIONS: With broader understanding of casemix adjustment and methods for analyzing small samples, quality data can be analysed and reported more accurately.

Data Interpretation, Statistical↗

Assessing data quality for decision support--emphasis on secondary analysis.

In secondary analysis, the use of available data makes it possible for the researcher to bypass the most time-consuming and costly steps in the research process. However, there are some noteworthy pitfalls and problems in working with existing data, especially the uncertainty of data quality. If the integrity and quality of data are not assured, statistical analysis of the data will not be reliable, no matter what statistical procedure is used. If the analysis is not reliable, the information used for decision support will not be accurate.

Data Collection↗

HGVbase: a human sequence variation database emphasizing data quality and a broad spectrum of data sources.

HGVbase (Human Genome Variation database; http://hgvbase.cgb.ki.se, formerly known as HGBASE) is an academic effort to provide a high quality and non-redundant database of available genomic variation data of all types, mostly comprising single nucleotide polymorphisms (SNPs). Records include neutral polymorphisms as well as disease-related mutations. Online search tools facilitate data interrogation by sequence similarity and keyword queries, and searching by genome coordinates is now being implemented. Downloads are freely available in XML, Fasta, SRS, SQL and tagged-text file formats. Each entry is presented in the context of its surrounding sequence and many records are related to neighboring human genes and affected features therein. Population allele frequencies are included wherever available. Thorough semi-automated data checking ensures internal consistency and addresses common errors in the source information. To keep pace with recent growth in the field, we have developed tools for fully automated annotation. All variants have been uniquely mapped to the draft genome sequence and are referenced to positions in EMBL/GenBank files. Data utility is enhanced by provision of genotyping assays and functional predictions. Recent data structure extensions allow the capture of haplotype and genotype information, and a new initiative (along with BiSC and HUGO-MDI) aims to create a central repository for the broad collection of clinical mutations and associated disease phenotypes of interest.

Base Sequence↗

Data quality assurance and quality control measures in large multicenter stroke trials: the African-American Antiplatelet Stroke Prevention Study experience.

Data quality assurance and quality control are critical to the effective conduct of a clinical trial. In the present commentary, we discuss our experience in a large, multicenter stroke trial. In addition to standard data quality control techniques, we have developed novel methods to enhance the entire process. Central to our methods is the use of clinical monitors who are trained in the techniques of data monitoring.

Journal Article↗

One fish, two fish, we QC fish: controlling data quality among more than 50 organizations over a four-year period.

EPA is conducting a National Study of Chemical Residues in Lake Fish Tissue. The study involves five analytical laboratories, multiple sampling teams from each of the 47 participating states, several tribes, all 10 EPA Regions and several EPA program offices, with input from other federal agencies. To fulfill study objectives, state and tribal sampling teams are voluntarily collecting predator and bottom-dwelling fish from approximately 500 randomly selected lakes over a 4-year period. The fish will be analyzed for more than 300 pollutants. The long-term nature of the study, combined with the large number of participants, created several QA challenges: (1) controlling variability among sampling activities performed by different sampling teams from more than 50 organizations over a 4-year period; (2) controlling variability in lab processes over a 4-year period; (3) generating results that will meet the primary study objectives for use by OW statisticians; (4) generating results that will meet the undefined needs of more than 50 participating organizations; and (5) devising a system for evaluating and defining data quality and for reporting data quality assessments concurrently with the data to ensure that assessment efforts are streamlined and that assessments are consistent among organizations. This paper describes the QA program employed for the study and presents an interim assessment of the program's effectiveness.

Animals↗

Census data quality--a user's view.

"This paper presents the perspective of a major user of both decennial and economic [U.S.] census data. It illustrates how these data are used as a framework for commercial marketing research surveys that measure television audiences and sales of consumer goods through retail stores, drawing on Nielsen's own experience in data collection and evaluation. It reviews Nielsen's analyses of census data quality based, in part, on actual field evaluation of census results. Finally, it suggests ways that data quality might be evaluated and improved to enhance the usefulness of these census programs."

Americas↗

Data quality in population-based cancer registration: an assessment of the Merseyside and Cheshire Cancer Registry.

Merseyside and Cheshire Cancer Registry (MCCR) data quality was assessed by applying literature-based measures to 27,942 cases diagnosed in 1990 and 1991. Registrations after death (n = 8535) were also audited (n = 917) to estimate death certificate only (DCO) case accuracy and the proportion of registrations notified by death certificate (DC). Ascertainment appeared to be high from the registration/mortality ratio for lung [1.01:1] and to be low from capture-recapture estimates (59.4%), varying significantly with site from oesophagus [92.2% (95% CI 88.5-95.9)] to breast [47.5 (95% CI 41.8-53.2)]. The estimated DC-dependent proportion was 20% (5601 out of 27 942) with successful traceback in 3533 out of 5601 (63.1%) cases. DCO flagging (2497 out of 27,942, 8.9%) overestimated true DCO cases (2068 out of 27,942, 7.4%). The proportion of cases of unknown primary site was low (1.5%), varying significantly with age [0-4.2%, (95% CI 2.5-5.9)] and district [0.8% (95% CI 0.3-1.3) to 2.2% (95% CI 1.8-2.6)]. The median diagnosis to registration interval appeared to be good (10 weeks), varying significantly with site (P < 0.0001), age (P < 0.0001) and district (P < 0.0001). The proportion with a verified diagnosis was 77.3%, varying significantly with site [lung 55.2% (95% CI 53.7-56.7) to cervix 96.9% (95% CI 96.3-97.5)], age [45.2% (95% CI 40.9-49.5) to 97.5% (95% CI 96.4-98.6)] and district [71.8% (95% CI 69.9-73.8) to 82.5% (95% CI 80.7-84.3)]. The DCO percentages varied similarly by site [non-melanoma skin 0.4% (95% CI 0.2-0.6) to lung 22.6% CI (95% 19.9-25.3)], age [0.7(95% CI 0.1-1.4) to 23.0 (95% CI 19.4-26.6)] and district [6.9% (95% CI 5.7-8.1) to 13.9% (95% CI 12.9-15.0)]. MCCR data quality varied with age, site and district - inviting action - and apparently compares favourably with elsewhere, although deficiencies in published data hampered definitive assessment. Putting quality assurance into practice identified shortcomings in the scope, definition and application of existing measures, and absent standards impeded interpretation. Cancer registry quality assurance should henceforward be within an explicit framework of agreed and standardized measures.

Confidence Intervals↗

Data quality in computerized patient records. Analysis of a haematology biopsy report database.

This paper addresses the problem of data quality in electronic patient records using a computerized haematology biopsy report system as an example. Physicians extracted five parameters from a traditional free text cytology report and encoded these parameters thus producing a computer processable report. The parameters were 1) the organ biopsied, 2) quality of specimen, 3) cytological diagnosis including 4) a modifier code for the main diagnosis code (i.e. status post chemotherapy, Y-code) and 5) an additional key describing the degree of remission obtained after chemotherapy of acute leukemias. From the various steps involved in generating the electronic record we selected two critical ones: encoding of free text terms by physician staff; entering of the coded terms into a computer by lab staff. We analyzed the rates of correct, incorrect and missing codes for each of the five parameters. Our findings indicate that in this model of an electronic patient record: 1) there is significant inaccuracy of physicians during the process of encoding the free text report with error rates between 3.2 and 28% and omission rates up to 64%. 2) lab staff entering these coded data into the computer introduce additional errors (0-7.8%) but rarely miss correctly encoded data (0-0.9%). 3) introducing a revised coding system data quality improved significantly (p < or = 0.001) with a fivefold increase of correct and a 75% reduction of missing codes. 4) the clinical relevance of the diagnoses encoded as perceived by clinicians is a significant factor affecting error and omission rates.(ABSTRACT TRUNCATED AT 250 WORDS)

Adult↗

From figures to facts: data quality managers emerge as knowledge leaders.

It is no surprise that data quality managers are being recognized as knowledge leaders--they turn patient care data into valuable information for research, quality reviews, and planning. This article explores how this position is emerging in the HIM field and becoming a key player in quality care.

Benchmarking↗

Study finds purchasers aren't using providers' outcomes data, quality report cards.

Are your data collection efforts for naught? A recent study shows that employers and purchasers overlook provider outcomes data and quality report cards when choosing health plans and providers--opting instead for customer satisfaction, HEDIS, and cost data. Here are the details on the study, plus advice from one of the study's authors.

Consumer Behavior↗

The time interval between death and next-of-kin contact and its effects on response rates and data quality.

The relation of the interval of time between death and next-of-kin contact to outcome variables including response rates and data quality was examined in a nationally representative sample of 17,713 deaths of persons 25 years of age or older that occurred in the United States in 1986. For most of the outcome variables examined, the length of time had little effect, although there was a small decrease in the response rate and a small increase in the refusal rate for contact 40 or more weeks after death. The small decrease in the response rate and small increase in the refusal rate for the longest interval examined held for most decedent background characteristics examined (age, race, cause of death, and type of informant.) Authorizations to contact health care facilities signed by the respondents decreased slightly as the interval increased. The rate of returned mailed questionnaires passing quality and consistency edits increased slightly with time since death. Substantive responses (versus blanks, don't knows, etc.) decreased as time since death increased. Certain questions such as those on income and birth control pill use showed a decrease in response with time since death. Overall, the effects of longer time intervals between death and next-of-kin contact were less than expected on response rates and data quality, although our findings may reflect the high overall response rate, 90.5%, leaving little opportunity for significant areas of nonresponse.

Adult↗

Building data quality into clinical trials.

Meaningful data begin with the collection process. Pharmaceutical companies are using several different strategies in clinical trials to ensure the highest quality of data. This article will examine these approaches, with an emphasis on case report form development through database release.

Abstracting and Indexing↗

Is shorter always better? Relative importance of questionnaire length and cognitive ease on response rates and data quality for two dietary questionnaires.

In this study, the authors sought to determine the effects of length and clarity on response rates and data quality for two food frequency questionnaires (FFQs): the newly developed 36-page Diet History Questionnaire (DHQ), designed to be cognitively easier for respondents, and a 16-page FFQ developed earlier for the Prostate, Lung, Colorectal, and Ovarian (PLCO) Cancer Screening Trial. The PLCO Trial is a 23-year randomized controlled clinical trial begun in 1992. The sample for this substudy, which was conducted from January to April of 1998, consisted of 900 control and 450 screened PLCO participants aged 55-74 years. Controls received either the DHQ or the PLCO FFQ by mail. Screenees, who had previously completed the PLCO FFQ at baseline, were administered the DHQ. Among controls, the response rate for both FFQs was 82%. Average amounts of time needed by controls to complete the DHQ and the PLCO FFQ were 68 minutes and 39 minutes, respectively. Percentages of missing or uninterpretable responses were similar between instruments for questions on frequency of intake but were approximately 3 and 9 percentage points lower (p < or = 0.001) in the DHQ for questions on portion size and use of vitamin/mineral supplements, respectively. Among screenees, response rates for the DHQ and the PLCO FFQ were 84% and 89%, respectively, and analyses of questions on portion size and supplement use showed few differences. These data indicated that the shorter FFQ was not better from the perspective of response rate and data quality, and that clarity and ease of administration may compensate for questionnaire length.

Aged↗

Data quality of administratively collected hospital discharge data for liver cirrhosis epidemiology.

We estimated the validity, i.e., whether the diagnostic criteria were fulfilled for the patients registered with the diagnosis of liver cirrhosis in a Danish hospital discharge registry, and the completeness, i.e., whether all patients with liver cirrhosis were included in the registry. Information in the regional hospital discharge registry in the Country of Aarhus, Denmark was compared with hospital records and information in a pathology registry. 85.4% of the patients registered with a diagnosis of liver cirrhosis fulfilled the diagnostic criteria for the diagnosis (validity). 93.2% of the patients registered with biopsy proven liver cirrhosis in the pathology registry were found in the discharge registry (completeness) with a diagnosis of liver cirrhosis. The hospital discharge registry showed relatively few misclassifications and the Danish National Registry of Patients (NRP), which is based on the regional registries, may provide a unique study base for future research.

Biopsy↗

The Primary Care Assessment Survey: tests of data quality and measurement performance.

OBJECTIVES: The authors examine the data quality and measurement performance of the Primary Care Assessment Survey (PCAS), a patient-completed questionnaire that operationalizes formal definitions of primary care, including the definition recently proposed by the Institute of Medicine Committee on the Future of Primary Care. METHODS: The PCAS measures seven domains of care through 11 summary scales: accessibility (organizational, financial), continuity (longitudinal, visit-based), comprehensiveness (contextual knowledge of patient, preventive counseling), integration, clinical interaction (clinician-patient communication, thoroughness of physical examinations), interpersonal treatment, and trust. Data from a study of Massachusetts state employees (n = 6094) were used to evaluate key measurement properties of the 11 PCAS scales. Analyses were performed on the combined population and for each of the 16 subgroups defined according to sociodemographic and health characteristics. RESULTS: The 11 PCAS scales demonstrated consistently strong measurement characteristics across all subgroups of this adult population. Tests of scaling assumptions for summated rating scales were well satisfied by all Likert-scaled measures. Assessment of data completeness, scale score dispersion characteristics, and inter-scale correlations provide strong evidence for the soundness of all scales, and for the value of separately measuring and interpreting these concepts. CONCLUSIONS: With public and private sector policies increasingly emphasizing the importance of primary care, the need for tools to evaluate and improve primary care performance is clear. The PCAS has excellent measurement properties, and performs consistently well across varied segments of the adult population. Widespread application of an assessment methodology, such as the PCAS, will afford an empiric basis through which to measure, monitor, and continuously improve primary care.

Adult↗