PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Performance benchmarking”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

An introduction to benchmarking in healthcare.

Benchmarking--the process of establishing a standard of excellence and comparing a business function or activity, a product, or an enterprise as a whole with that standard--will be used increasingly by healthcare institutions to reduce expenses and simultaneously improve product and service quality. As a component of total quality management, benchmarking is a continuous process by which an organization can measure and compare its own processes with those of organizations that are leaders in a particular area. Benchmarking should be viewed as a part of quality management programs, not as a replacement. There are four kinds of benchmarking: internal, competitive, functional and generic. With internal benchmarking, functions within an organization are compared with each other. Competitive benchmarking partners do business in the same market and provide a direct comparison of products or services. Functional and generic benchmarking are performed with organizations which may have a specific similar function, such as payroll or purchasing, but which otherwise are in a different business. Benchmarking must be a team process because the outcome will involve changing current practices, with effects felt throughout the organization. The team should include members who have subject knowledge; communications and computer proficiency; skills as facilitators and outside contacts; and sponsorship of senior management. Benchmarking requires quantitative measurement of the subject. The process or activity that you are attempting to benchmark will determine the types of measurements used. Benchmarking metrics usually can be classified in one of four categories: productivity, quality, time and cost-related.

Outcome and Process Assessment, Health Care↗

Physician resource profiling enhances utilization management.

Physician resource profiling, the analysis of a physician's resource consumption, enhances performance uniformity and efficiency and assists in utilization management. Developing reliable profiles requires the shared participation of an organization's finance personnel and physicians. Selecting or developing benchmarks for performance comparisons, assessing the integrity of the organizational data used, and testing the developed profiles all should be completed before the physician resource profiling can be used to support decision making.

Benchmarking↗

An integrated care pathway for the last two days of life: Wales-wide benchmarking in palliative care.

Functional benchmarking assesses performance and practice across a broad range of settings and carries the potential to effect change in practice. An integrated care pathway (ICP) can assist in the benchmarking process, defining desired outcomes for specific patient groups over a designated time frame. Any variations to the agreed course of care are documented using the 'variance sheet'. This article describes the Wales-wide implementation of an ICP for the last two days of life. The project has enabled an ongoing centralized collection and analysis of variance sheets, which reflect the care of the dying patient in four different care settings crossing the voluntary and statutory sectors. Initial analysis of the first 500 variance sheets to be generated by the ICP for the last two days of life indicates that the management of pain, agitation, excess respiratory secretions and mouth care may be problematic. The same problems were experienced across acute, hospice, specialist inpatient units and community care. Closing the audit cycle involves incorporating the information from the variance analysis into clinical practice.

Analysis of Variance↗

Same-day surgery has new benchmarking option. Patient survey offers national comparisons.

A new survey instrument developed by the Picker Institute in Boston uses highly specific questions to target same-day surgery successes and failures from the patients' perspectives. Hospitals and surgery centers can compare their performance with benchmarks from the academic institutions that are members of the University HealthSystem Consortium in Oak Brook, IL. Surveys will identify which specialty areas produce the most and least satisfied patients.

Ambulatory Surgical Procedures↗

Multicenter study of oxygen-insensitive handheld glucose point-of-care testing in critical care/hospital/ambulatory patients in the United States and Canada.

OBJECTIVES: Existing handheld glucose meters are glucose oxidase (GO)-based. Oxygen side reactions can introduce oxygen dependency, increase potential error, and limit clinical use. Our primary objectives were to: a) introduce a new glucose dehydrogenase (GD)-based electrochemical biosensor for point-of-care testing; b) determine the oxygen-sensitivity of GO- and GD-based electrochemical biosensor test strips; and c) evaluate the clinical performance of the new GD-based glucose meter system in critical care/hospital/ambulatory patients. DESIGN: Multicenter study sites compared glucose levels determined with GD-based biosensors to glucose levels determined in whole blood with a perchloric acid deproteinization hexokinase reference method. One site also studied GO-based biosensors and venous plasma glucose measured with a chemistry analyzer. Biosensor test strips were used with a handheld glucose monitoring system. Bench and clinical oxygen sensitivity, hematocrit effect, and precision were evaluated. SETTING: The study was performed at eight U.S. medical centers and one Canadian medical center. PATIENTS: There were 1,248 patients. RESULTS: The GO-based biosensor was oxygen-sensitive. The new GD-based biosensor was oxygen-insensitive. GD-based biosensor performance was acceptable: 2,104 (96.1%) of 2,189 glucose meter measurements were within +/-15 mg/dL (+/-0.83 mmol/L) for glucose levels of < or = 100 mg/dL (< or = 5.55 mmol/L) or within +/-15% for glucose levels of > 100 mg/dL, compared with the whole-blood reference method results. With the GD-based biosensor, the percentages of glucose measurements that were not within the error tolerance were comparable for different specimen types and clinical groups. Bracket predictive values were acceptable for glucose levels used in therapeutic management. CONCLUSIONS: The performance of GD-based, oxygen-insensitive, handheld glucose testing was technically suitable for arterial specimens in critical care patients, cord blood and heelstick specimens in neonates, and capillary and venous specimens in other patients. Multicenter findings benchmark the performance of bedside glucose testing devices. With the new +/-15 mg/dL --> 100 mg/dL --> +/-15% accuracy criterion, point-of-care systems for handheld glucose testing should score 95% (or better), as compared with the recommended reference method. Physiologic changes, preanalytical factors, confounding variables, and treatment goals must be taken into consideration when interpreting glucose results, especially in critically ill patients, for whom arterial blood glucose measurements will reflect systemic glucose levels.

Adult↗

Trends in acute myocardial infarction management: use of the National Registry ofMyocardial Infarction in quality improvement.

Cardiovascular disease, including acute myocardial infarction (AMI), is the leading cause of death in the United States and was the primary disease category among hospital discharges in 1996. Efforts to improve hospital care of patients with AMI should be measured and assessed routinely for appropriateness of care and improvement of medical staff performance. The National Registry of Myocardial Infarction (NRMI), an observational Phase IV study, has enrolled > 1 million AMI patients since 1990, and is now in its third phase. NRMI 3 collects patient data and facilitates the measurement of improvement in care and outcomes, while allowing participating institutions to benchmark their performance against national, state, and like-hospital data. Three measures from NRMI 3 are accepted for the Joint Commission on Accreditation of Healthcare Organizations' ORYX initiative: (1) aspirin use within 24 hours of AMI diagnosis; (2) door-to-drug time for fibrinolysis; and (3) no initial reperfusion strategy given to eligible patients.

Anti-Inflammatory Agents, Non-Steroidal↗

Benchmarking and quality in residential and nursing homes: lessons from the US.

BACKGROUND: Performance measurement and benchmarking are common concerns in the delivery of long term care. It is common to measure the performance of providers and to publicly report these data. This paper examines selected technical challenges facing those who design, implement and disseminate health care quality performance measures. METHOD: Review of the application of measures of performance in the US nursing home sector. RESULTS: Using examples drawn from the skilled nursing home arena, problems ranging from data reliability and validity, the multi-dimensional nature of quality measures and selection bias as well as differential measurement abilities are discussed. CONCLUSIONS: Benchmarking of performance is an inherently complex issue. However, to ensure that such comparisons are both fair and valid requires measures to be more technically sophisticated and sensitive to real changes attributable to changes in care.

Benchmarking↗

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score&#xa0;=&#xa0;0.67-0.90) in genus diversity and showed a high correlation (rSpearman&#xa0;=&#xa0;0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics↗

Value decision-making: staff benchmarking.

Benchmarking is becoming a more important management tool--especially for setting staff levels. MGMA data, from Cost Surveys and Physician Compensation and Productivity Surveys, can help group managers set realistic goals. However, if taken simply at face value, the data may not provide adequate specificity; it may not convey the quality and value staff provide a particular organization. This paper show how to use MGMA data to perform staff benchmarking.

Benchmarking↗

Nonparametric regression applied to quantitative structure-activity relationships

Several nonparametric regressors have been applied to modeling quantitative structure-activity relationship (QSAR) data. The simplest regressor, the Nadaraya-Watson, was assessed in a genuine multivariate setting. Other regressors, the local linear and the shifted Nadaraya-Watson, were implemented within additive models--a computationally more expedient approach, better suited for low-density designs. Performances were benchmarked against the nonlinear method of smoothing splines. A linear reference point was provided by multilinear regression (MLR). Variable selection was explored using systematic combinations of different variables and combinations of principal components. For the data set examined, 47 inhibitors of dopamine beta-hydroxylase, the additive nonparametric regressors have greater predictive accuracy (as measured by the mean absolute error of the predictions or the Pearson correlation in cross-validation trails) than MLR. The use of principal components did not improve the performance of the nonparametric regressors over use of the original descriptors, since the original descriptors are not strongly correlated. It remains to be seen if the nonparametric regressors can be successfully coupled with better variable selection and dimensionality reduction in the context of high-dimensional QSARs.

Journal Article↗

An audit of error rates in a UK district hospital transfusion laboratory.

We have audited the error rates of our transfusion laboratory and compared these with error rates reported in the transfusion literature. Error rates were calculated using workload data from the department. The majority of errors that were detected were preanalytical and related to inadequate or incomplete data provided on the sample or request form. These errors were all corrected prior to any further action being taken on that request. The main analytical errors were transcription errors in entering patient identification information into the laboratory computer by transfusion staff together with the incorrect performance of blood group testing. For postanalytical errors the main errors were failure of nursing staff to follow procedures for the collection of blood components prior to transfusion. There were no serious consequences identified of the errors detected in this study. It was difficult to compare these results with those published in the literature in view of the different methodologies that have been reported when error rates have been determined. A standard method should be developed in the UK for calculating error rates so that laboratories can benchmark their performance against comparable organizations.

Blood Grouping and Crossmatching↗

Application of non-parametric regression to quantitative structure-activity relationships.

Several non-parametric regressors have been applied to modelling quantitative structure-activity relationship (QSAR) data. Performances were benchmarked against multilinear regression and the nonlinear method of smoothing splines. Variable selection was explored through systematic combinations of different variables and combinations of principal components. For the training set examined--539 inhibitors of the tyrosine kinase, Syk--the best two-descriptor model had a 5-fold cross-validated q2 of 0.43. This was generated by a multi-variate Nadaraya-Watson kernel estimator. A subsequent, independent, test set of 371 similar chemical entities showed the model had some predictive power. Other approaches did not perform as well. A modest increase in predictive ability can be achieved with three descriptors, but the resulting model is less easy to visualise. We conclude that non-parametric regression offers a potentially powerful approach to identifying predictive, low-dimensional QSARs.

Databases, Factual↗

Re-evaluating genetic algorithm performance under coordinate rotation of benchmark functions. A survey of some theoretical and practical aspects of genetic algorithms.

In recent years, genetic algorithms (GAs) have become increasingly robust and easy to use. Current knowledge and many successful experiments suggest that the application of GAs is not limited to easy-to-optimize unimodal functions. Several results and GA theory give the impression that GAs easily escape from millions of local optima and reliably converge to a single global optimum. The theoretical analysis presented in this paper shows that most of the widely-used test functions have n independent parameters and that, when optimizing such functions, many GAs scale with an O(n ln n) complexity. Furthermore, it is shown that the current design of GAs and its parameter settings are optimal with respect to independent parameters. Both analysis and results show that a rotation of the coordinate system causes a severe performance loss to GAs that use a small mutation rate. In case of a rotation, the GA's complexity can increase up to O(nn) = O(exp(n ln n)). Future work should find new GA designs that solve this performance loss. As long as these problems have not been solved, the application of GAs will be limited to the optimization of easy-to-optimize functions.

Algorithms↗

TargetQC: A targeted quality control framework for clinical genomic testing.

Reliable genetic testing depends on accurate assessment of sequencing quality in clinically relevant genomic regions that directly influence variant interpretation. We developed TargetQC, a flexible quality control framework that supports user-defined gene sets, coverage thresholds, and variant sets for evaluating sequencing performance across exome sequencing (ES) and genome sequencing (GS) platforms. TargetQC assesses exon and gene coverage, identifies regions meeting predefined coverage thresholds, evaluates variant detection accuracy, and measures sequencing quality at pathogenic variant sites. We applied TargetQC to the reference sample NA12878 and 665 clinical samples across five ES platforms and one GS platform. ES-VendorB and ES-VendorE achieved the most complete coverage of OMIM coding regions in NA12878, whereas ES-VendorD and ES-VendorE showed the highest coverage compliance in clinical samples. ES-VendorB and GS demonstrated the highest variant detection accuracy. TargetQC provides a practical framework for benchmarking sequencing performance and informing platform selection in clinical genomics.

exome sequencing↗

Using risk assessment to evaluate adverse selection under capitated contracts.

Most provider organizations rely on health plan or market-based information about capitation rates, per member per month costs, and utilization trends to benchmark their performance. However, these statistics can be misleading because of differences in enrollee mix and contracting terms across provider organizations. This article describes the limitations of health plan contract provisions in protecting against adverse selection. It describes various actuarial and statistical data sources for evaluation of adverse selection. The article then presents various approaches to risk adjustment on population basis and their use in quantifying adverse selection for health plan contract negotiations.

Actuarial Analysis↗

Scalable medium-density genotyping platforms for cultivar identification, pedigree authentication, marker-assisted and genomic selection, and other applications in strawberry.

A broad spectrum of high-density genotyping approaches, including single-nucleotide polymorphism (SNP) arrays, genotyping-by-sequencing, and whole-genome reduced-representation sequencing, have been shown to perform well in strawberry (Fragaria &#xd7; ananassa), despite the inherent complexity of the octoploid genome. While these approaches are effective, their routine deployment in breeding programs can be constrained by cost, computational requirements, and workflow complexity. In parallel, many breeding programs continue to rely on locus-specific assays for marker-assisted selection, resulting in fragmented and inefficient genotyping strategies. Here, we describe medium-density amplicon-based genotyping platforms for strawberry designed to provide cost-effective, turnkey solutions that integrate markers used for marker-assisted selection with genome-wide markers suitable for genomic prediction in a single laboratory assay. These platforms were developed by targeting 1,650 or 4,811 target SNPs via amplicon sequencing, and are interoperable with existing high-density genotyping resources, including a widely used 50K SNP array, thereby facilitating data integration across platforms. We benchmarked their performance relative to the 50K SNP array across breeding-relevant applications, including identity and purity testing, pedigree authentication, marker-assisted selection, and genomic selection, and further evaluated the feasibility of genotype imputation to enhance genome-wide information content. Across analyses, the 1,650- and 4,811-amplicon platforms produced results comparable to higher-density platforms while substantially reducing genotyping cost and analytical overhead. This work demonstrates that targeted amplicon-based genotyping can support efficient, scalable, and integrated genome-informed breeding, enabling the routine application of both marker-assisted and genomic selection within strawberry breeding workflows. Open-source R workflows are provided to support streamlined analyses in breeding contexts.

Fragaria↗

Validating the performance of the mammary sentinel lymph node team.

BACKGROUND AND OBJECTIVES: The mammary sentinel lymph node (SLN) procedure has the potential to improve the accuracy and lower the morbidity of axillary staging in breast cancer patients, but results are closely linked to experience and can vary widely between institutions. Standardized performance measures need to be established in order to optimize the transition to SLN biopsy only. METHODS: Performance data were prospectively collected for the first 156 mammary SLN procedures performed by three surgeons in our institution. RESULTS: Seventy-five cases were required to achieve an SLN visualization rate of > 80% on preoperative lymphoscintigraphy. The SLN visualization rate was 90% for the last 52 cases. Two surgeons required 25 cases before consistently achieving a > or = 90% SLN identification rate in the operating room and one required 15 cases. The metastasis detection rate increased from 22% for the first 52 cases to 31% for the last 52 cases. The false negative rate for the procedure was 5%. CONCLUSIONS: The following performance criteria and benchmarks are suggested for validating the performance of the SLN team: (1) SLN visualization rate on preoperative lymphoscintigraphy > or = 80%, (2) SLN identification rate in the operating room > or = 90%, (3) False negative rate for the procedure 5%. Thirty procedures per surgeon were sufficient to achieve these benchmarks in our group.

Axilla↗

Automated Classification of Lymphoma Subtypes From Histopathological Images Using a U-Net Deep Learning Model: Comparative Evaluation Study.

BACKGROUND: Accurate classification and grading of lymphoma subtypes are essential for treatment planning. Traditional diagnostic methods face challenges of subjectivity and inefficiency, highlighting the need for automated solutions based on deep learning techniques. OBJECTIVE: This study aimed to investigate the application of deep learning technology, specifically the U-Net model, in classifying and grading lymphoma subtypes to enhance diagnostic precision and efficiency. METHODS: In this study, the U-Net model was used as the primary tool for image segmentation integrated with attention mechanisms and residual networks for feature extraction and classification. A total of 620 high-quality histopathological images representing 3 major lymphoma subtypes were collected from The Cancer Genome Atlas and the Cancer Imaging Archive. All images underwent standardized preprocessing, including Gaussian filtering for noise reduction, histogram equalization, and normalization. Data augmentation techniques such as rotation, flipping, and scaling were applied to improve the model's generalization capability. The dataset was divided into training (70%), validation (15%), and test (15%) subsets. Five-fold cross-validation was used to assess model robustness. Performance was benchmarked against mainstream convolutional neural network architectures, including fully convolutional network, SegNet, and DeepLabv3+. RESULTS: The U-Net model achieved high segmentation accuracy, effectively delineating lesion regions and improving the quality of input for classification and grading. The incorporation of attention mechanisms further improved the model's ability to extract key features, whereas the residual structure of the residual network enhanced classification accuracy for complex images. In the test set (N=1250), the proposed fusion model achieved an accuracy of 92% (1150/1250), a sensitivity of 91.04% (1138/1250), a specificity of 89.04% (1113/1250), and an F1-score of 90% (1125/1250) for the classification of the 3 lymphoma subtypes, with an area under the receiver operating characteristic curve of 0.95 (95% CI 0.93-0.97). The high sensitivity and specificity of the model indicate strong clinical applicability, particularly as an assistive diagnostic tool. CONCLUSIONS: Deep learning techniques based on the U-Net architecture offer considerable advantages in the automated classification and grading of lymphoma subtypes. The proposed model significantly improved diagnostic accuracy and accelerated pathological evaluation, providing efficient and precise support for clinical decision-making. Future work may focus on enhancing model robustness through integration with advanced algorithms and validating performance across multicenter clinical datasets. The model also holds promise for deployment in digital pathology platforms and artificial intelligence-assisted diagnostic workflows, improving screening efficiency and promoting consistency in pathological classification.

Humans↗