PubMed HealthSearch

SEARCH · PubMed Health

Results for “data quality”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The time interval between death and next-of-kin contact and its effects on response rates and data quality.

The relation of the interval of time between death and next-of-kin contact to outcome variables including response rates and data quality was examined in a nationally representative sample of 17,713 deaths of persons 25 years of age or older that occurred in the United States in 1986. For most of the outcome variables examined, the length of time had little effect, although there was a small decrease in the response rate and a small increase in the refusal rate for contact 40 or more weeks after death. The small decrease in the response rate and small increase in the refusal rate for the longest interval examined held for most decedent background characteristics examined (age, race, cause of death, and type of informant.) Authorizations to contact health care facilities signed by the respondents decreased slightly as the interval increased. The rate of returned mailed questionnaires passing quality and consistency edits increased slightly with time since death. Substantive responses (versus blanks, don't knows, etc.) decreased as time since death increased. Certain questions such as those on income and birth control pill use showed a decrease in response with time since death. Overall, the effects of longer time intervals between death and next-of-kin contact were less than expected on response rates and data quality, although our findings may reflect the high overall response rate, 90.5%, leaving little opportunity for significant areas of nonresponse.

Adult

An investigation into the distribution of radial immunodiffusion quality control data.

Quality control data from routine radial immunodiffusion assays for IgG, IgA, IgM, C3, C4 and alpha1-antitrypsin were tested by the Kolmogorov--Smirnov procedure for gaussian distribution. All but alpha1-antitrypsin were nongaussian in type. Further analysis of these date by plotting on log-normal probability paper showed them to have a log-normal distribution. Treatment of the date by either gaussian or nonparametric statistical methods produced little difference in confidence limits. It does not appear necessary to use nonparametric methods to calculate confidence limits from quality control data for the procedures studied.

Complement C3

Establishing a national pediatric stem cell transplantation registry in Iran addressing implementation and data quality challenges.

The Iranian Pediatric Hematopoietic Stem Cell Transplantation Registry (IPED-HSCT) was established to enhance data collection, improve patient outcomes, and support clinical research in pediatric hematopoietic stem cell transplantation. This study aimed to assess the feasibility and reliability of implementing a standardized registry in pediatric settings. A community-based participatory study was conducted across three pediatric HSCT centers in Iran. The registry development involved a multi-phase approach, including pilot testing and the implementation of a web-based system. Data were collected from fifty pediatric patients who underwent HSCT for both malignant and non-malignant conditions, with a focus on data completeness and user satisfaction. Statistical analyses were performed using IBM SPSS Statistics. The registry achieved a data completeness rate exceeding 90%, with a participant demographic of 31 males (62%) and 19 females (38%). Rigorous quality control measures and real-time validation rules were implemented, enhancing data reliability. User feedback indicated high satisfaction with the platform's design and training sessions. Challenges included variations in long-term follow-up data collection across centers. The IPED-HSCT Registry demonstrates that establishing a robust pediatric HSCT registry is feasible even in resource-limited settings. Its innovative features offer a scalable model for similar initiatives in developing countries. Future research should focus on ensuring long-term sustainability and fostering international collaborations to improve pediatric HSCT outcomes globally. not applicable.

Humans

Improvements in data quality in the USRDS database: determining treatment modalities.

Past USRDS estimates of the prevalent ESRD population have exceeded the counts reported by the HCFA Annual Facility Surveys. One expects the USRDS estimates to be lower because the Facility Surveys include non-Medicare patients generally not in the USRDS database. The methodology for determining the treatment histories of patients has been modified to define lost to follow-up periods as periods of at least one year during which the patient has no dialysis data and does not have a functioning transplant. Patients are not counted as prevalent when they are in such a lost to follow-up period. This change brings the USRDS year end prevalent counts down to about 94 percent of the Facility Survey counts of total dialysis patients and slightly over the Facility Survey counts of Medicare dialysis patients. This change raises prevalent mortality rates by about seven percent over the rates reported in the 1991 USRDS Annual Data Report. We expect to make further refinements in this methodology.

Data Collection

Evaluating Wearable Devices for Remote Monitoring in Psychosis: Pilot Study Nested Within the CONNECT Cohort Study.

BACKGROUND: Digital remote monitoring technologies, including smartphones and wearables, offer promising avenues for early detection of psychosis relapse. However, selecting devices that are acceptable to participants and produce high-quality data remains challenging. OBJECTIVE: The aim of this nested pilot study was to assess the acceptability and data quality of 3 commercially available wearable devices in people with psychosis recruited to the CONNECT cohort study. METHODS: Participants recruited to the CONNECT study before July 31, 2024, were included in the pilot study and selected 1 of 3 wearable devices: a Fitbit Charge 5, Samsung Galaxy Watch 5, or Apple Watch SE. Baseline demographics were compared between device groups. Acceptability of devices to participants was assessed through a Wearable Device Satisfaction Questionnaire after 3 months of use, with the proportion of positive responses to each question calculated and compared. Data completeness was also assessed by calculating the number (and percentage) of valid days of step count, heart rate, and sleep data, and comparing between groups. Data quality was assessed through summarizing the amount of troubleshooting required, additional metrics available from the wearables, and continuity of data completeness by calculating the proportion of participants with at least 3 days of heart rate data per week for the first 20 weeks of follow-up. Predefined criteria were used to determine the next steps for the wider CONNECT study: if one device was superior, this would be selected; if none were found to be superior and the Fitbit was found to be noninferior, then Fitbit would be retained. RESULTS: Of the first 107 participants recruited to CONNECT, 105 were included in the pilot study evaluation. The Samsung Galaxy Watch was selected most frequently by participants (46/105, 43.8%), followed by the Apple Watch (27/105, 25.7%), and Fitbit Charge (23/105, 21.9%). Differences in participant demographics were observed across device groups. Self-reported acceptability after use did not differ substantially between devices. However, in terms of data completeness, the median proportion of valid heart rate data days was significantly lower for Samsung Galaxy (median 31.2%, IQR 8.5%-46.0%) compared to Fitbit (median 80.1%, IQR 26.7%-95.0%; P=.003) and Apple Watch (median 49.3%, IQR 21.5%-86.0%; P=.02). There was no significant difference between Fitbit and Apple Watch. Similar patterns were observed for step count and sleep data. The Samsung Galaxy Watch required more frequent troubleshooting for data flow issues and lacked additional physiological metrics, available from the other devices. CONCLUSIONS: Due to comparatively lower data quality and technical performance, the Samsung Galaxy Watch was discontinued for use in the subsequent phase of the CONNECT study. The study highlights the importance of incorporating nested evaluations of devices in long-term research.

Humans

The positive known association design: a quality assurance method for occupational health surveillance data.

Quality control must be an integral component of an occupational health surveillance program. The positive known association design offers the occupational health physician a method to test, on a population basis (ie, high periodic medical surveillance examination participation rates by the employees), the quality of periodic medical surveillance data. Several well-established biological associations were evaluated and observed in this study, including a dramatic relation between white blood cell counts and smoking. We highly recommend that the positive known association design be incorporated in the quality assurance procedures of occupational health surveillance programs.

Adult

Quality of data in the Manchester orthopaedic database.

OBJECTIVE: To determine the completeness and accuracy of data in a computerised clinical information system (Manchester orthopaedic database) in comparison with the data available through the Hospital Activity Analysis. DESIGN: Retrospective review of case notes, computer data, and Hospital Activity Analysis data. SETTING: Orthopaedic unit in a district general hospital in Manchester. SUBJECTS: 200 random patient records distributed through the period of use of the computer system (1 October 1988 to 31 March 1990) and 121 records for random admissions between 1 April 1989 and 31 March 1990, 71 of which were included in the previous sample. MAIN OUTCOME MEASURES: Conformity of the computer record key words and Hospital Activity Analysis codes to an ideal key word record and ideal code record drawn up by one investigator from the clinical notes; overall quality (completeness times accuracy). RESULTS: Overall completeness of the data in the orthopaedic database was 62% and the accuracy was 96%. Completeness improved after feedback to doctors on the use of key words in regular audit meetings. Completeness was higher in inpatient than outpatient records (69.9% v 53.7%, p less than 0.001) and when a new key word was required compared with missing and incorrect key words (both p less than 0.001). Completeness was lower when the key word was required of a senior registrar (p less than 0.05). Accuracy was not significantly different. The completeness of Hospital Activity Analysis data was 90.5% and accuracy 69.5%. Thus the overall data quality was similar in both systems. CONCLUSIONS: Even in a system designed for simple and efficient data capture, compliance by users was poor. Accuracy was high, suggesting that users understood the principles of data entry. Completeness of data capture can be improved by providing feedback to users on use of the system and performance. Improvements in future versions of the software should improve performance.

Abstracting and Indexing

A systematic comparison of three structure determination methods from NMR data: dependence upon quality and quantity of data.

We have systematically examined how the quality of NMR protein structures depends on (1) the number of NOE distance constraints, (2) their assumed precision, (3) the method of structure calculation and (4) the size of the protein. The test sets of distance constraints have been derived from the crystal structures of crambin (5 kDa) and staphylococcal nuclease (17 kDa). Three methods of structure calculation have been compared: Distance Geometry (DGEOM), Restrained Molecular Dynamics (XPLOR) and the Double Iterated Kalman Filter (DIKF). All three methods can reproduce the general features of the starting structure under all conditions tested. In many instances the apparent precision of the calculated structure (as measured by the RMS dispersion from the average) is greater than its accuracy (as measured by the RMS deviation of the average structure from the starting crystal structure). The global RMS deviations from the reference structures decrease exponentially as the number of constraints is increased, and after using about 30% of all potential constraints, the errors asymptotically approach a limiting value. Increasing the assumed precision of the constraints has the same qualitative effect as increasing the number of constraints. For comparable numbers of constraints/residue, the precision of the calculated structure is less for the larger than for the smaller protein, regardless of the method of calculation. The accuracy of the average structure calculated by Restrained Molecular Dynamics is greater than that of structures obtained by purely geometric methods (DGEOM and DIKF).

Magnetic Resonance Spectroscopy

ATAC-seq in Emerging Model Organisms: Challenges and Strategies.

The Assay for Transposase-Accessible Chromatin with sequencing (ATAC-seq) is a versatile and widely utilized method for identifying potential regulatory regions, such as promoters and enhancers, within a genome. ATAC-seq has been successfully applied to a wide range of established and emerging model organisms. However, implementing this method in emerging model systems, such as arthropods, can be challenging due to several factors that influence data quality. These factors include the availability of a sufficient amount and quality of tissue or cells, the need for species- and tissue-specific protocol optimization, the completeness and accuracy of the reference genome, and the quality of the genome annotation. In this article, we emphasize the key steps in the ATAC-seq protocol that, based on our experience, have the greatest impact on data quality when adapting this method for emerging model organisms. Specifically, we discuss the importance of nuclei isolation, the incubation conditions of the Tn5 transposase, and PCR amplification of the library. Furthermore, we outline essential quality checkpoints during the bioinformatic analysis of ATAC-seq data to assist in assessing data integrity and consistency. Given that many emerging model organisms may not be readily available in laboratory cultures, we also emphasize the importance of evaluating how different preservation methods affect ATAC-seq data quality. Based on examples in one spider and one ant species, we demonstrate that replication and thorough quality controls at all steps of the protocol and data analysis are essential to assess the usability of ATAC-seq data. Our data highlights the importance of isolating the right number of intact nuclei, as well as ensuring optimal amplification conditions during library preparation to obtain good-quality sequence data for downstream analyses. We recommend using fresh tissue samples if possible because we show that direct cryopreservation of the tissue may affect chromatin integrity. This effect could be avoided or reduced by preserving the homogenate in cell culture medium. Overall, we explain the ATAC-seq protocol and downstream analyses in detail and give step-by-step advice to researchers who are new to the field and want to implement this method. With careful planning and validation, ATAC-seq can reveal the regulatory landscape of a genome and aid in identifying elements that govern gene expression.

Animals

Evaluation of surgical services in a large university-affiliated VA hospital: use of an in-house-generated quality assurance data base.

In this era of occurrence screening, increased documentation of surgical resident supervision, and overall efforts to alter quality of patient care through increased documentation, an in-house quality assurance data base can greatly facilitate these activities. To establish a data base, 6241 operative procedures over a 15-month interval were logged into a StatView II Database using the patient's name and record number, diagnosis, and operative information available from the standard operation report worksheet. Additional information, including the name of the supervising staff and level of supervision, had also been recorded concurrently on this form. Each month, after this demographic information has been entered into the data base for each surgical service, it is reviewed, ensuring 100% surveillance of all cases. The subsequent morbidity and mortality (M & M) information is then added to the data base. Using the contingency table option of the StatView II Database, quarterly and annual reviews have been done to study individual services, types of cases, staff attending, individual supervision, levels of staff involvement, and incidence of complications. Our Surgical Service has reported an overall morbidity of 306 cases (4.9%) and mortality of 71 (1.14%). Average numerically coded levels of resident supervision have been compiled with regard to individual services and staff surgeons. This data base has proved helpful in the recredentialling process of our surgical staff, resident logs, identification of necessary reviews of certain case types, and formation of a computerized operative log.

Computers

Performance of techniques for measurement of therapeutic drugs in serum. A comparison based on external quality assessment data.

Ten assay techniques were compared using measurements of a range of 15 drugs spiked in freeze-dried samples of serum reported to the Heathcontrol External Quality Assessment Scheme between November 1988 and January 1991. Three measures of performance were studied: frequency of outliers greater than 3 standard deviations from the sample mean, the coefficient of variation (CV) of sample measurements, and the difference of the sample mean from the spike value. The most consistently precise technique was polarisation fluoroimmunoassay (PFIA). It was in the group of techniques producing significantly fewer outliers and lower CVs than other techniques for all its target analytes. However, a specific interaction with the animal serum used as sample matrix resulted in significant negative bias in PFIA measurements of carbamazepine. Other immunoassay techniques and high-performance liquid chromatography also performed well for a range of analytes, in most cases giving less than 6% of outliers with CVs of less than 13% and less than 5% bias. The least satisfactory techniques were nephelometry and gas-liquid chromatography with derivatisation, which for several analytes gave significantly more outliers and higher CV values than other techniques. In samples containing carbamazepine-10, 11-epoxide, immunoassay measurements of carbamazepine showed cross-reactivity with the epoxide metabolite of between 7 and 15%.

Carbamazepine

Comprehensive evaluation of new sequencer T20 and well-established T7 with 507 human samples.

The DNBSEQ-T20×2 (T20) sequencer, developed by MGI Tech, enables cost-effective human whole-genome sequencing (WGS) at 30× coverage for less than $100 per genome. Here, we evaluate the sequencing performance and data quality of the T20 platform by benchmarking it against the established DNBSEQ-T7 (T7) sequencer using 507 samples derived from blood (N = 75), stool (N = 242), and saliva (N = 190). The T20 exhibited lower sequencing quality metrics compared with the T7, with Q20 scores of 95.76%-95.83% and Q30 scores of 87.25%-87.40%, compared with 97.81%-97.93% and 93.26%-93.60%, respectively, for T7 data. Quality differences were more evident toward the end of reads, and PCR-free libraries sequenced on the T20 showed similar reductions in quality scores. The median empirical base error rate estimated from 102 ZymoBIOMICS samples was 0.33%. The T20 demonstrated comparable coverage uniformity to the T7 and showed high concordance in microbiome composition analysis, with a median Bray-Curtis dissimilarity of 0.02. Variant calling performance was highly consistent between the two platforms. Among variants with non-missing genotype calls on both platforms, 94.92% of SNPs and 87.20% of InDels showed concordant genotypes between T20 and T7. Overall, the T20 delivers reliable sequencing accuracy and reproducibility for large-scale genomic and microbiome studies, providing a cost-effective alternative for high-throughput sequencing applications.

Metagenomics