PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Performance benchmarking”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Resolution-dependent self-supervised transfer in chest radiograph classification.

BACKGROUND: Self-supervised learning (SSL) has improved visual representation learning, but its value in chest radiography remains uncertain. DINOv3 extends earlier SSL models through Gram-anchored self-distillation and explicit high-resolution adaptation. Whether these changes improve transfer learning for chest radiograph classification has not been established. METHODS: We benchmarked DINOv3 against DINOv2 and supervised ImageNet initialization across seven chest radiograph datasets comprising 816,183 radiographs from pediatric and adult cohorts. ViT-B/16 and ConvNeXt-B were evaluated under full fine-tuning at 224 × 224 and 512 × 512 pixels, with targeted 1024 × 1024 experiments on three cohorts. Additional analyses examined parameter-efficient adaptation, synthetic label corruption, external validation, frozen 7B features, and computational efficiency. The primary outcome was the mean area under the receiver operating characteristic curve across labels. RESULTS: In adult cohorts, DINOv3 did not consistently outperform DINOv2 at 224 × 224 pixels, but became the strongest initialization at 512 × 512 pixels, especially with ConvNeXt-B. Gains were greatest for small focal and boundary-dependent abnormalities, whereas large-structure findings changed little. The pediatric cohort showed no significant benefit from DINOv3, higher resolution, or backbone choice. Scaling to 1024 × 1024 rarely improved performance and markedly increased computational cost. ConvNeXt-B remained superior to ViT-B/16 under both full and parameter-efficient adaptation. External validation preserved the 512 × 512 DINOv3 advantage, whereas synthetic label corruption showed that this benefit should not be interpreted simply as superior noise robustness. Frozen DINOv3-7B features underperformed relative to fully adapted 86 to 89M-parameter backbones. CONCLUSIONS: For adult chest radiograph classification, DINOv3 provides its most reliable benefit at 512 × 512 pixels, particularly with ConvNeXt-B. Fully adapted mid-sized models at 512 × 512 pixels provided the best performance-cost trade-off in our benchmark.

Journal Article↗

Comparing biofilm models for a single species biofilm system.

A benchmark problem was defined to evaluate the performance of different mathematical biofilm models. The biofilm consisted of heterotrophic bacteria degrading organic substrate and oxygen. Mathematical models tested ranged from simple analytical to multidimensional numerical models. For simple and more or less flat biofilms it was shown that analytical biofilm models provide very similar results compared to more complex numerical solutions. When considering a heterogeneous biofilm morphology it was shown that the effect of an increased external mass transfer resistance was much more significant compared to the effect of an increased surface area inside the biofilm.

Bacteria↗

Mortality prediction using SAPS II: an update for French intensive care units.

INTRODUCTION: The standardized mortality ratio (SMR) is commonly used for benchmarking intensive care units (ICUs). Available mortality prediction models are outdated and must be adapted to current populations of interest. The objective of this study was to improve the Simplified Acute Physiology Score (SAPS) II for mortality prediction in ICUs, thereby improving SMR estimates. METHOD: A retrospective data base study was conducted in patients hospitalized in 106 French ICUs between 1 January 1998 and 31 December 1999. A total of 77,490 evaluable admissions were split into a training set and a validation set. Calibration and discrimination were determined for the original SAPS II, a customized SAPS II and an expanded SAPS II developed in the training set by adding six admission variables: age, sex, length of pre-ICU hospital stay, patient location before ICU, clinical category and whether drug overdose was present. The training set was used for internal validation and the validation set for external validation. RESULTS: With the original SAPS II calibration was poor, with marked underestimation of observed mortality, whereas discrimination was good (area under the receiver operating characteristic curve 0.858). Customization improved calibration but had poor uniformity of fit; discrimination was unchanged. The expanded SAPS II exhibited good calibration, good uniformity of fit and better discrimination (area under the receiver operating characteristic curve 0.879). The SMR in the validation set was 1.007 (confidence interval 0.985-1.028). Some ICUs had better and others worse performance with the expanded SAPS II than with the customized SAPS II. CONCLUSION: The original SAPS II model did not perform sufficiently well to be useful for benchmarking in France. Customization improved the statistical qualities of the model but gave poor uniformity of fit. Adding simple variables to create an expanded SAPS II model led to better calibration, discrimination and uniformity of fit, producing a tool suitable for benchmarking.

Adult↗

An iterative refinement algorithm for consistency based multiple structural alignment methods.

MOTIVATION: Multiple STructural Alignment (MSTA) provides valuable information for solving problems such as fold recognition. The consistency-based approach tries to find conflict-free subsets of alignments from a pre-computed all-to-all Pairwise Alignment Library (PAL). If large proportions of conflicts exist in the library, consistency can be hard to get. On the other hand, multiple structural superposition has been used in many MSTA methods to refine alignments. However, multiple structural superposition is dependent on alignments, and a superposition generated based on erroneous alignments is not guaranteed to be the optimal superposition. Correcting errors after making errors is not as good as avoiding errors from the beginning. Hence it is important to refine the pairwise library to reduce the number of conflicts before any consistency-based assembly. RESULTS: We present an algorithm, Iterative Refinement of Induced Structural alignment (IRIS), to refine the PAL. A new measurement for the consistency of a library is also proposed. Experiments show that our algorithm can greatly improve T-COFFEE performance for less consistent pairwise alignment libraries. The final multiple alignment outperforms most state-of-the-art MSTA algorithms at assembling 15 transglycosidases. Results on three other benchmarks showed that the algorithm consistently improves multiple alignment performance. AVAILABILITY: The C++ code of the algorithm is available upon request.

Algorithms↗

[Benchmarks for surgical gynecology: results of the German Society of Gynecology and Obstetrics Quality Assurance Study].

Profiling of performance and quality in gynecological surgery is discussed. Unfortunately, most report cards miss valid clinical benchmarks. Within the German study on quality assurance in gynecological surgery we explored whether indicators of quality were suitable as clinical benchmarks. Using a factor analytic approach, we reduced the number of indicators and obtained in a set of 13 indicators of clinical quality. On the basis of the study data on post operative infections we show that these indicators are suitable as clinical benchmarks: the clinical benchmarks are able to make health care quality transparent and demonstrate opportunities for improvement of the processes of gynecological care.

Benchmarking↗

Cardiovascular benchmarking saves hospital nearly $897,000.

When Chattanooga, TN-based Erlanger Medical Center wanted to implement care paths for cardiac patients, its benchmarking team found it had to treat its cardiovascular surgeons as one unit rather than as individual practitioners. Best practices emerged as a combination of internal benchmarking and study of care paths from top performing hospitals in the region and nationally. When incorporated in Erlanger's new care plans, these practices netted savings of $896,000 through reduced stays, utilization, and standardization of supplies and processes.

Benchmarking↗

Benchmarking in emergency health systems.

This paper discusses the role of benchmarking as a component of quality management. It describes the historical background of benchmarking, its competitive origin and the requirement in today's health environment for a more collaborative approach. The classical 'functional and generic' types of benchmarking are discussed with a suggestion to adopt a different terminology that describes the purpose and practicalities of benchmarking. Benchmarking is not without risks. The consequence of inappropriate focus and the need for a balanced overview of process is explored. The competition that is intrinsic to benchmarking is questioned and the negative impact it may have on improvement strategies in poorly performing organizations is recognized. The difficulty in achieving cross-organizational validity in benchmarking is emphasized, as is the need to scrutinize benchmarking measures. The cost effectiveness of benchmarking projects is questioned and the concept of 'best value, best practice' in an environment of fixed resources is examined.

Benchmarking↗

Study provides comparative targets to help reduce costs of surgical procedures.

Data Benchmarks: Cost and LOS benchmarks can help your organization target and reduce costs of surgical procedures. Setting up clinical pathways to improve quality and reduce costs requires good benchmarking data to zero in on the appropriate clinical services or procedures. This month's Data Benchmarks offers good comparative data on the most cost-intensive surgical procedures performed at U.S. hospitals.

Ancillary Services, Hospital↗

Gaining competitive advantage in personal dosimetry services through ISO 9001 certification.

This paper discusses the advantage of certification process in the quality assurance of individual dose monitoring in Malaysia. The demand by customers and the regulatory authority for a higher degree of quality service requires a switch in emphasis from a technically focused quality assurance program to a comprehensive quality management for service provision. Achieving the ISO 9001:2000 certification by an accredited third party demonstrates acceptable recognition and documents the fact that the methods used are capable of generating results that satisfy the performance criteria of the certification program. It also offers a proof of the commitment to quality and, as a benchmark, allows measurement of the progress for continual improvement of service performance.

Body Burden↗

Connectionist-based Dempster-Shafer evidential reasoning for data fusion.

Dempster-Shafer evidence theory (DSET) is a popular paradigm for dealing with uncertainty and imprecision. Its corresponding evidential reasoning framework is theoretically attractive. However, there are outstanding issues that hinder its use in real-life applications. Two prominent issues in this regard are 1) the issue of basic probability assignments (masses) and 2) the issue of dependence among information sources. This paper attempts to deal with these issues by utilizing neural networks in the context of pattern classification application. First, a multilayer perceptron neural network with the mean squared error as a cost function is implemented to calculate, for each information source, posteriori probabilities for all classes. Second, an evidence structure construction scheme is developed for transferring the estimated posteriori probabilities to a set of masses along with the corresponding focal elements, from a Bayesian decision point of view. Third, a network realization of the Dempster-Shafer evidential reasoning is designed and analyzed, and it is further extended to a DSET-based neural network, referred to as DSETNN, to manipulate the evidence structures. In order to tackle the issue of dependence between sources, DSETNN is tuned for optimal performance through a supervised learning process. To demonstrate the effectiveness of the proposed approach, we apply it to three benchmark pattern classification problems. Experiments reveal that the DSETNN out-performs DSET and provide encouraging results in terms of classification accuracy and the speed of learning convergence.

Algorithms↗

Interlaboratory comparison of immunohistochemical testing for HER2: results of the 2004 and 2005 College of American Pathologists HER2 Immunohistochemistry Tissue Microarray Survey.

CONTEXT: Correct assessment of human epidermal growth factor receptor 2 (HER2) status is essential in managing patients with invasive breast carcinoma, but few data are available on the accuracy of laboratories performing HER2 testing by immunohistochemistry (IHC). OBJECTIVE: To review the results of the 2004 and 2005 College of American Pathologists HER2 Immunohistochemistry Tissue Microarray Survey. DESIGN: The HER2 survey is designed for laboratories performing immunohistochemical staining and interpretation for HER2. The survey uses tissue microarrays, each consisting of ten 3-mm tissue cores obtained from different invasive breast carcinomas. All cases are also analyzed by fluorescence in situ hybridization. Participants receive 8 tissue microarrays (80 cases) with instructions to perform immunostaining for HER2 using the laboratory's standard procedures. The laboratory interprets the stained slides and returns results to the College of American Pathologists for analysis. In 2004 and 2005, a core was considered "graded" when at least 90% of laboratories agreed on the result--negative (0, 1+) versus positive (2+, 3+). This interlaboratory comparison survey included 102 laboratories in 2004 and 141 laboratories in 2005. RESULTS: Of the 160 cases in both surveys, 111 (69%) achieved 90% consensus (graded). All 43 graded cores scored as IHC-positive were fluorescence in situ hybridization-positive, whereas all but 3 of the 68 IHC-negative graded cores were fluorescence in situ hybridization-negative. Ninety-seven (95%) of 102 laboratories in 2004 and 129 (91%) of 141 laboratories in 2005 correctly scored at least 90% of the graded cores. CONCLUSION: Performance among laboratories performing HER2 IHC in this tissue microarray-based survey was excellent. Cores found to be IHC-positive or IHC-negative by participant consensus can be used as validated benchmarks for interlaboratory comparison, allowing laboratories to assess their performance and determine if improvements are needed.

Breast Neoplasms↗

Guidelines for internal peer review in the cardiac catheterization laboratory. Laboratory Performance Standards Committee, Society for Cardiac Angiography and Interventions.

The Laboratory Performance Standards Committee of the Society for Cardiac Angiography and Interventions has proposed guidelines for establishing an internal peer review program in the cardiac catheterization laboratory. The first step is to establish a committee and a data base. This data base should include quality indicators that reflect: physician qualifications, outcomes of procedures, and processes of care. The outcomes must be risk-adjusted to account for the variable severity of illness. Data should be collected by catheterization laboratory personnel and entered into a laboratory-specific computerized data base. These data must be analyzed and organized into profiles that reflect the quality of care. Based on this information, the Committee would institute the following interventions to improve physician performance: education, clinical practice standardization, feedback and benchmarking, professional interaction, incentives, decision-support systems, and administrative interventions. The legal aspects of peer review are reviewed briefly.

Cardiac Catheterization↗

An adaptive and iterative algorithm for refining multiple sequence alignment.

Multiple sequence alignment is a basic tool in computational genomics. The art of multiple sequence alignment is about placing gaps. This paper presents a heuristic algorithm that improves multiple protein sequences alignment iteratively. A consistency-based objective function is used to evaluate the candidate moves. During the iterative optimization, well-aligned regions can be detected and kept intact. Columns of gaps will be inserted to assist the algorithm to escape from local optimal alignments. The algorithm has been evaluated using the BAliBASE benchmark alignment database. Results show that the performance of the algorithm does not depend on initial or seed alignments much. Given a perfect consistency library, the algorithm is able to produce alignments that are close to the global optimum. We demonstrate that the algorithm is able to refine alignments produced by other software, including ClustalW, SAGA and T-COFFEE. The program is available upon request.

Algorithms↗

A comparison of match-only algorithms for the analysis of Plasmodium falciparum oligonucleotide arrays.

This study is motivated by two data sets which employ a custom Plasmodium falciparum version of the Affymetrix GeneChip, containing only perfect match (PM) oligonucleotides. A PM-only chip cannot be analysed using the standard Affymetrix-supplied software. We compared the performance of three match-only algorithms on these data: the Match Only Integral Distribution (MOID) algorithm, Robust Multichip Analysis (RMA), and the Model Based Expression Index (MBEI). We validated the differential expression of several genes using quantitative reverse transcriptase-PCR. We also performed a comparison using two publicly available 'benchmarking' data sets: the Latin Square spike-in data set generated by Affymetrix, and the Gene Logic dilution series. Since we know what the true fold changes are in these special data sets, they are helpful for assessment of expression algorithms.

Algorithms↗

Variant harmonization critically determines polygenic score transferability for lipid traits in Samoan populations.

Dyslipidemia is a significant risk factor for cardiovascular disease (CVD), the leading cause of death in Samoa. Polygenic scores (PGSs) for lipid traits offer promise for improved CVD risk prediction; however, their performance in Pacific Islander populations-comprising only 0.002% of genome-wide association study (GWAS) participants as of 2024-remains unknown. We evaluated the transferability of multi-ancestry PGS for LDL cholesterol (LDL-C), HDL cholesterol (HDL-C), triglycerides (TGs), and total cholesterol (TC) in 4,342 Samoan adults across five cohorts spanning 1990-2010. PGSs from Graham et al. and Kanoni et al. multi-ancestry meta-analyses were harmonized with genome-wide imputed genotypes using a Samoan-specific reference panel, and performance was assessed via incremental R2 from linear mixed models with bootstrapped confidence intervals. HDL-C showed the highest performance (incremental R2 5.0%-15.0%), followed by TC (5.0%-10.7%), LDL-C (5.7%-8.6%), and TG (3.5%-7.0%). Critically, meaningful LDL-C performance was achieved only with the genome-wide PRS-CS score (99.6%-99.7% variant matching), while a curated pruning-and-thresholding score achieved ∼9% matching and near-zero performance. These findings establish systematic lipid PGS benchmarks in Samoans, demonstrating meaningful transferability when genome-wide variant coverage is ensured, and highlight variant harmonization as a critical precondition for PGS deployment in underrepresented populations.

Pacific Islanders↗

Adaptation of the CAS test system and synthetic sewage for biological nutrient removal. Part II: design and validation of test units.

A global increase in biological nutrient removal (BNR) applications in wastewater treatment and concern for potential effects of anthropogenic substances on BNR processes resulted in the adaptation of the Continuous Activated Sludge (CAS) laboratory test system (cf. guideline OECD 303A or ISO 11733). In this paper two novel systems are compared to the standard CAS unit: the Behrotest KLD4 and a University of Cape Town system (CAS-UCT). Both are 'single sludge' systems with an anoxic/aerobic and an anaerobic/anoxic/aerobic configuration, respectively. They both can simulate the essential processes of full-scale BNR installations. The units where fed with a specially designed synthetic sewage, Syntho (cf. Part I of this study), or its precursor BSR3 medium. The performance of the two new units was benchmarked against the standard CAS system in terms of carbon/nitrogen/phosphorus (C/N/P) removal, as well as primary biodegradation of the surfactants linear alkylbenzene sulfonate (LAS) and glucose amide (GA). Both systems allow to easily achieve stable excess N- and P-removal. Experimental C/N/P removal data compared closely with simulations obtained with the IAWQ Activated Sludge Model No. 2 (ASM2), and with full scale BNR plants with a similar configuration. In both units the effluent concentrations of the surfactants tested were significantly reduced in comparison to the standard CAS system (up to 50% less). No adverse effects on BNR were noted for the test surfactants dosed at 400 microg/l together with an overall surfactant background concentration in the feed of ca. 20 mg/l. The proposed systems hold potential to complement the standard CAS system for situations where advanced sewage treatment plants with BNR need to be simulated in the laboratory with minimum effort.

Anaerobiosis↗

Fuzzy least squares support vector machines for multiclass problems.

In least squares support vector machines (LS-SVMs), the optimal separating hyperplane is obtained by solving a set of linear equations instead of solving a quadratic programming problem. But since SVMs and LS-SVMs are formulated for two-class problems, unclassifiable regions exist when they are extended to multiclass problems. In this paper, we discuss fuzzy LS-SVMs that resolve unclassifiable regions for multiclass problems. We define a membership function in the direction perpendicular to the optimal separating hyperplane that separates a pair of classes. Using the minimum or average operation for these membership functions, we define a membership function for each class. Using some benchmark data sets, we show that recognition performance of fuzzy LS-SVMs with the minimum operator is comparable to that of fuzzy SVMs, but fuzzy LS-SVMs with the average operator showed inferior performance.

Fuzzy Logic↗