PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Performance benchmarking”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Validating the performance of the mammary sentinel lymph node team.

BACKGROUND AND OBJECTIVES: The mammary sentinel lymph node (SLN) procedure has the potential to improve the accuracy and lower the morbidity of axillary staging in breast cancer patients, but results are closely linked to experience and can vary widely between institutions. Standardized performance measures need to be established in order to optimize the transition to SLN biopsy only. METHODS: Performance data were prospectively collected for the first 156 mammary SLN procedures performed by three surgeons in our institution. RESULTS: Seventy-five cases were required to achieve an SLN visualization rate of > 80% on preoperative lymphoscintigraphy. The SLN visualization rate was 90% for the last 52 cases. Two surgeons required 25 cases before consistently achieving a > or = 90% SLN identification rate in the operating room and one required 15 cases. The metastasis detection rate increased from 22% for the first 52 cases to 31% for the last 52 cases. The false negative rate for the procedure was 5%. CONCLUSIONS: The following performance criteria and benchmarks are suggested for validating the performance of the SLN team: (1) SLN visualization rate on preoperative lymphoscintigraphy > or = 80%, (2) SLN identification rate in the operating room > or = 90%, (3) False negative rate for the procedure 5%. Thirty procedures per surgeon were sufficient to achieve these benchmarks in our group.

Axilla↗

Automated Classification of Lymphoma Subtypes From Histopathological Images Using a U-Net Deep Learning Model: Comparative Evaluation Study.

BACKGROUND: Accurate classification and grading of lymphoma subtypes are essential for treatment planning. Traditional diagnostic methods face challenges of subjectivity and inefficiency, highlighting the need for automated solutions based on deep learning techniques. OBJECTIVE: This study aimed to investigate the application of deep learning technology, specifically the U-Net model, in classifying and grading lymphoma subtypes to enhance diagnostic precision and efficiency. METHODS: In this study, the U-Net model was used as the primary tool for image segmentation integrated with attention mechanisms and residual networks for feature extraction and classification. A total of 620 high-quality histopathological images representing 3 major lymphoma subtypes were collected from The Cancer Genome Atlas and the Cancer Imaging Archive. All images underwent standardized preprocessing, including Gaussian filtering for noise reduction, histogram equalization, and normalization. Data augmentation techniques such as rotation, flipping, and scaling were applied to improve the model's generalization capability. The dataset was divided into training (70%), validation (15%), and test (15%) subsets. Five-fold cross-validation was used to assess model robustness. Performance was benchmarked against mainstream convolutional neural network architectures, including fully convolutional network, SegNet, and DeepLabv3+. RESULTS: The U-Net model achieved high segmentation accuracy, effectively delineating lesion regions and improving the quality of input for classification and grading. The incorporation of attention mechanisms further improved the model's ability to extract key features, whereas the residual structure of the residual network enhanced classification accuracy for complex images. In the test set (N=1250), the proposed fusion model achieved an accuracy of 92% (1150/1250), a sensitivity of 91.04% (1138/1250), a specificity of 89.04% (1113/1250), and an F1-score of 90% (1125/1250) for the classification of the 3 lymphoma subtypes, with an area under the receiver operating characteristic curve of 0.95 (95% CI 0.93-0.97). The high sensitivity and specificity of the model indicate strong clinical applicability, particularly as an assistive diagnostic tool. CONCLUSIONS: Deep learning techniques based on the U-Net architecture offer considerable advantages in the automated classification and grading of lymphoma subtypes. The proposed model significantly improved diagnostic accuracy and accelerated pathological evaluation, providing efficient and precise support for clinical decision-making. Future work may focus on enhancing model robustness through integration with advanced algorithms and validating performance across multicenter clinical datasets. The model also holds promise for deployment in digital pathology platforms and artificial intelligence-assisted diagnostic workflows, improving screening efficiency and promoting consistency in pathological classification.

Humans↗

Evaluation and development of potentially better practices to prevent chronic lung disease and reduce lung injury in neonates.

OBJECTIVE: Despite increased knowledge and improving technology, chronic lung disease (CLD) rates in extremely low birth weight infants have remained constant for 20 years. One reason for this is an ineffective translation of research-proven improvements into practice. The Neonatal Intensive Care Quality Improvement Collaborative Year 2000 (NIC/Q 2000) was created to provide participating nurseries the tools necessary to effect change. The objective of this study was to develop and implement a process that uses quality improvement techniques to collaboratively improve CLD rates. METHODS: Nine member hospitals of the NIC/Q 2000 collaborative formed a focus group aiming to decrease CLD rates. The focus group established goals and outcome measures, created a list of potentially better practices (PBPs) based on available literature, benchmarked and performed site visits, encouraged individual site implementation of PBPs, developed a database, and measured outcomes. RESULTS: The goal "decrease CLD rates in extremely low birth weight infants" was established. Nine PBPs were identified, and 57 PBPs were implemented by the 9 participating sites. Twelve site visits were conducted, and a 435-patient database of infants with a mean birth weight of 789 g was established. CONCLUSIONS: Collaborative use of quality improvement techniques resulted in creation of a logical, efficient, and effective process to improve CLD rates. Group creation of PBPs, based on literature review and reinforced with site visits, internal data analysis, and improved individual site outcomes, resulted in accelerated and effective change, unlikely to occur if attempted outside of the collaborative.

Benchmarking↗

[Precautions and limitations when using intrahospital mortality as indicator of quality of care].

INTRODUCTION: Inhospital mortality has been used as an outcome quality indicator in the USA and in England to compare and benchmark hospital performance. It is now also possible to measure this outcome indicator in Switzerland, but it is important to highlight limitations and precautions to its use. METHODS: We collected administrative data from acute care community hospitals in the Canton of Valais, Switzerland, for the year 2001. We assessed rates of global and disease specific inhospital mortality and calculated crude and adjusted relative risks of inhospital mortality, specific to each hospital. RESULTS: The crude rates of the global inhospital mortality varied from 1.25% to 1.80% between hospitals. The variation for disease specific mortality rates was low. After adjustment, differences between relative risks were almost never statistically significant. DISCUSSION: The use of inhospital mortality as an quality indicator, need to be done with cautions, in particular adjustment for the case-mix, exclusion of patients in palliative care and analysis of disease specific rates.

Benchmarking↗

In vivo dehydration of silicone hydrogel contact lenses.

PURPOSE: To benchmark the performance of new-generation silicone hydrogel contact lenses in terms of their in vivo hydration characteristics and to highlight the possible clinical ramifications of any changes observed. METHODS: Thirteen subjects (four men and nine women with a mean age of 24.8 +/- 5.0 years) wore a silicone hydrogel lens (PureVision, Bausch & Lomb, Rochester, NY) in one eye and a conventional hydrogel lens (ACUVUE 2, Johnson & Johnson Vision Care, Jacksonville, FL) in the other eye for 4 weeks on an extended-wear basis. A gravimetric method was used to determine lens water content and dehydration during the intended lifespan of the lenses. RESULTS: For the PureVision lens, the water content was 38.3% +/- 0.9%, 35.2% +/- 1.1%, and 35.3% +/- 1.7% at baseline and after 2 and 4 weeks of wear, respectively (F=28.4, P<0.0001). For the ACUVUE 2 lens, the water content was 58.1% +/- 0.6% and 52.1% +/- 1.3% at baseline and after 2 weeks of wear, respectively. Thus, after 2 weeks of wear, absolute dehydration was 2.8% +/- 1.8% and 6.0% +/- 1.3% for the PureVision and ACUVUE 2 lenses, respectively (t=6.8, P<0.0001). The mass of deposition was calculated to be 568 +/- 457 microg for the PureVision lens and 1,660 +/- 499 microg for the ACUVUE 2 lens (t=5.1, P=0.0003). CONCLUSIONS: The ACUVUE 2 lens underwent a greater degree of lens dehydration, causing a reduction in oxygen permeability (3.6 barrer), and deposition after 2 weeks of extended wear. The loss of water from the PureVision lens was paradoxically associated with a 6.0-barrer increase in oxygen permeability after 4 weeks of extended wear.

Adult↗

Importance of the postdischarge interval in assessing major adverse clinical event rates following percutaneous coronary intervention.

In-hospital major adverse clinical event (MACE) rates after percutaneous coronary intervention serve as benchmarks of performance. However, accelerated clinical pathways, decreased lengths of stay, and potential delayed effects of percutaneous coronary intervention may result in an underestimation of this traditional measurement of outcome. Records from patients in the first 3 waves of the National Heart, Lung, and Blood Institute's Dynamic Registry (n = 6,676) were reviewed for rates of composite in-hospital MACEs (death, myocardial infarction, and any repeat target vessel revascularization) and postdischarge MACEs (death, myocardial infarction, repeat hospitalization, and repeat target vessel revascularization) through 30 days. Rates for each composite MACE were compared across waves to assess changes over time. Predictors of each MACE category were identified using multivariate analysis. In-hospital MACE decreased significantly (5.4% of wave 1, 4.9% of wave 2, 3.1% of wave 3, p <0.001), whereas stent implantation increased significantly (67.5% of wave 1, 79.1% of wave 2, 86.2% of wave 3, p <0.001). Postdischarge MACE through 30 days remained unchanged (5.1% of wave 1, 5.1% of wave 2, 4.8% of wave 3, p = 0.6). Mean length of stay decreased (2.7 days for wave 1, 2.2 days for wave 3, p <0.001). Disparate clinical, procedural, and angiographic factors were associated with each MACE. Postdischarge MACE rates through 30 days comprise a significant and unchanging fraction of overall procedurally related MACE rates despite improving in-hospital outcomes. Most postdischarge events derive from pathology related to the controlled vessel. A 30-day MACE rate may serve as a more comprehensive measurement of procedural outcome.

Aged↗

Top-performing physician groups focus on innovative cost control.

DATA BENCHMARKS: Top-performing physician group practices reveal why they are head and shoulders above their peers. A new study released by the Medical Group Management Association offers benchmark data that shows how leading group practices in various specialties are faring compared with their colleagues nationwide.

Benchmarking↗

Development of an analytical method for the determination of antidepressants in water samples by capillary electrophoresis with electrospray ionization mass spectrometric detection.

A method for the quantitative determination of major antidepressants in aqueous matrices by CE using ESI-MS is presented. Several aqueous, nonaquoeus, and mixed aqueous/organic solvent BGEs including inorganic and organic acids were investigated with respect to their suitability for the separation of the selected analytes. Finally, due to the necessity to employ MS detection if the developed method should be suitable also for environmental samples, only MS-compatible electrolytes were taken into account. Based on this fact optimum results were obtained with a system consisting of 1.5 M formic acid and 50 mM ammonium formate in ACN/water (85/15). Linear calibration plots could be obtained for all solutes over a concentration range of almost two orders of magnitude, and the LODs achieved were in the range of 3-6 microg/L for trazodone and 39-43 microg/L for sertraline with the TOF instrument and the single quadrupole instrument in the SIM mode, respectively. This fact allowed the assumption that the presented method can be regarded as suitable for the determination of antidepressants even in the trace amounts commonly present in environmental samples. Spiking of river water and sewage plant effluent extracts with the selected solutes showed that no interferences from the matrix usually found in such samples can be expected. Finally the quantitative determination of the seven antidepressants in environmental samples was used to benchmark the performance of CZE coupled to a single quadrupole MS and a TOF-MS.

Antidepressive Agents↗

SSALN: an alignment algorithm using structure-dependent substitution matrices and gap penalties learned from structurally aligned protein pairs.

In template-based modeling of protein structures, the generation of the alignment between the target and the template is a critical step that significantly affects the accuracy of the final model. This paper proposes an alignment algorithm SSALN that learns substitution matrices and position-specific gap penalties from a database of structurally aligned protein pairs. In addition to the amino acid sequence information, secondary structure and solvent accessibility information of a position are used to derive substitution scores and position-specific gap penalties. In a test set of CASP5 targets, SSALN outperforms sequence alignment methods such as a Smith-Waterman algorithm with BLOSUM50 and PSI_BLAST. SSALN also generates better alignments than PSI_BLAST in the CASP6 test set. LOOPP server prediction based on an SSALN alignment is ranked the best for target T0280_1 in CASP6. SSALN is also compared with several threading methods and sequence alignment methods on the ProSup benchmark. SSALN has the highest alignment accuracy among the methods compared. On the Fischer's benchmark, SSALN performs better than CLUSTALW and GenTHREADER, and generates more alignments with accuracy >50%, >60% or >70% than FUGUE, but fewer alignments with accuracy >80% than FUGUE. All the supplemental materials can be found at http://www.cs.cornell.edu/ approximately jianq/research.htm.

Algorithms↗

A branch and bound algorithm for protein structure refinement from sparse NMR data sets.

We describe new methods for predicting protein tertiary structures to low resolution given the specification of secondary structure and a limited set of long-range NMR distance constraints. The NMR data sets are derived from a realistic protocol involving completely deuterated 15N and 13C-labeled samples. A global optimization method, based upon a modification of the alphaBB (branch and bound) algorithm of Floudas and co-workers, is employed to minimize an objective function combining the NMR distance restraints with a residue-based protein folding potential containing hydrophobicity, excluded volume, and van der Waals interactions. To assess the efficacy of the new methodology, results are compared with benchmark calculations performed via the X-PLOR program of Brünger and co-workers using standard distance geometry/molecular dynamics (DGMD) calculations. Seven mixed alpha/beta proteins are examined, up to a size of 183 residues, which our methods are able to treat with a relatively modest computational effort, considering the size of the conformational space. In all cases, our new approach provides substantial improvement in root-mean-square deviation from the native structure over the DGMD results; in many cases, the DGMD results are qualitatively in error, whereas the new method uniformly produces high quality low-resolution structures. The DGMD structures, for example, are systematically non-compact, which probably results from the lack of a hydrophobic term in the X-PLOR energy function. These results are highly encouraging as to the possibility of developing computational/NMR protocols for accelerating structure determination in larger proteins, where data sets are often underconstrained.

Algorithms↗

Progress in numerical modelling of the Cl influence on gamma-ray spectra from an n-gamma logging tool, by using the improved ENDF data for radiative capture.

Quality of the numerical modelling (MCNP code) of the spectrometric neutron-gamma benchmark experiment, performed at the Polish Calibration Station BGW in Zielona Gora for quantification of the main rock elements: Si, Ca, Fe and H, is considered. Elemental concentrations obtained from the measurements and simulations, for the rock models with water-filled boreholes, are in good agreement. For chlorine present in the borehole, the quality of the numerical reproducibility of the measured elemental concentrations depends on the cross section library used for the Cl(n,gamma)Cl reaction. The standard evaluated nuclear data library ENDF/B-VI Release 2 supplies imperfect data for photon production from thermal neutron capture in Cl. The improved cross sections for Cl(n,gamma)Cl are included in the ENDF/B-VI Release 8 library. Superiority of this new compilation over the previous one is shown in the paper. The accuracies for the Si, Ca and Fe determination have been improved by about 36%, 19.9% and 21.4%, respectively, when the ENDF/B-VI Release 8 library has been used for Cl.

Journal Article↗

Deep-probe metal-clad waveguide biosensors.

Two types of metal-clad waveguide biosensors, so-called dip-type and peak-type, are analyzed and tested. Their performances are benchmarked against the well-known surface-plasmon resonance biosensor, showing improved probe characteristics for adlayer thicknesses above 150-200 nm. The dip-type metal-clad waveguide sensor is shown to be the best all-round alternative to the surface-plasmon resonance biosensor. Both metal-clad waveguides are tested experimentally for cell detection, showing a detection limit of 8-9 cells/mm2.

Biosensing Techniques↗

Utilizing data grid architecture for the backup and recovery of clinical image data.

Grid Computing represents the latest and most exciting technology to evolve from the familiar realm of parallel, peer-to-peer and client-server models. However, there has been limited investigation into the impact of this emerging technology in medical imaging and informatics. In particular, PACS technology, an established clinical image repository system, while having matured significantly during the past ten years, still remains weak in the area of clinical image data backup. Current solutions are expensive or time consuming and the technology is far from foolproof. Many large-scale PACS archive systems still encounter downtime for hours or days, which has the critical effect of crippling daily clinical operations. In this paper, a review of current backup solutions will be presented along with a brief introduction to grid technology. Finally, research and development utilizing the grid architecture for the recovery of clinical image data, in particular, PACS image data, will be presented. The focus of this paper is centered on applying a grid computing architecture to a DICOM environment since DICOM has become the standard for clinical image data and PACS utilizes this standard. A federation of PACS can be created allowing a failed PACS archive to recover its image data from others in the federation in a seamless fashion. The design reflects the five-layer architecture of grid computing: Fabric, Resource, Connectivity, Collective, and Application Layers. The testbed Data Grid is composed of one research laboratory and two clinical sites. The Globus 3.0 Toolkit (Co-developed by the Argonne National Laboratory and Information Sciences Institute, USC) for developing the core and user level middleware is utilized to achieve grid connectivity. The successful implementation and evaluation of utilizing data grid architecture for clinical PACS data backup and recovery will provide an understanding of the methodology for using Data Grid in clinical image data backup for PACS, as well as establishment of benchmarks for performance from future grid technology improvements. In addition, the testbed can serve as a road map for expanded research into large enterprise and federation level data grids to guarantee CA (Continuous Availability, 99.999% up time) in a variety of medical data archiving, retrieval, and distribution scenarios.

Diagnostic Imaging↗

Deterministic projection by growing cell structure networks for visualization of high-dimensionality datasets.

Recent advances in clinical proteomics data acquisition have led to the generation of datasets of high complexity and dimensionality. We present here a visualization method for high-dimensionality datasets that makes use of neuronal vectors of a trained growing cell structure (GCS) network for the projection of data points onto two dimensions. The use of a GCS network enables the generation of the projection matrix deterministically rather than randomly as in random projection. Three datasets were used to benchmark the performance and to demonstrate the use of this deterministic projection approach in real-life scientific applications. Comparisons are made to an existing self-organizing map projection method and random projection. The results suggest that deterministic projection outperforms existing methods and is suitable for the visualization of datasets of very high dimensionality.

Algorithms↗

Rapid decision threshold modulation by reward rate in a neural network.

Optimal performance in two-alternative, free response decision-making tasks can be achieved by the drift-diffusion model of decision making--which can be implemented in a neural network--as long as the threshold parameter of that model can be adapted to different task conditions. Evidence exists that people seek to maximize reward in such tasks by modulating response thresholds. However, few models have been proposed for threshold adaptation, and none have been implemented using neurally plausible mechanisms. Here we propose a neural network that adapts thresholds in order to maximize reward rate. The model makes predictions regarding optimal performance and provides a benchmark against which actual performance can be compared, as well as testable predictions about the way in which reward rate may be encoded by neural mechanisms.

Algorithms↗

Mining gene expression data using a novel approach based on hidden Markov models.

In this work we have developed a new framework for microarray gene expression data analysis. This framework is based on hidden Markov models. We have benchmarked the performance of this probability model-based clustering algorithm on several gene expression datasets for which external evaluation criteria were available. The results showed that this approach could produce clusters of quality comparable to two prevalent clustering algorithms, but with the major advantage of determining the number of clusters. We have also applied this algorithm to analyze published data of yeast cell cycle gene expression and found it able to successfully dig out biologically meaningful gene groups. In addition, this algorithm can also find correlation between different functional groups and distinguish between function genes and regulation genes, which is helpful to construct a network describing particular biological associations. Currently, this method is limited to time series data. Supplementary materials are available at http://www.bioinfo.tsinghua.edu.cn/~rich/hmmgep_supp/.

Algorithms↗

A novel learning algorithm which improves the partial fault tolerance of multilayer neural networks.

The paper deals with the problem of fault tolerance in a multilayer perceptron network. Although it already possesses a reasonable fault tolerance capability, it may be insufficient in particularly critical applications. Studies carried out by the authors have shown that the traditional backpropagation learning algorithm may entail the presence of a certain number of weights with a much higher absolute value than the others. Further studies have shown that faults in these weights is the main cause of deterioration in the performance of the neural network. In other words, the main cause of incorrect network functioning on the occurrence of a fault is the non-uniform distribution of absolute values of weights in each layer. The paper proposes a learning algorithm which updates the weights, distributing their absolute values as uniformly as possible in each layer. Tests performed on benchmark test sets have shown the considerable increase in fault tolerance obtainable with the proposed approach as compared with the traditional backpropagation algorithm and with some of the most efficient fault tolerance approaches to be found in literature.

Journal Article↗

Matrix-assisted laser desorption/ionization time-of-flight mass spectrometry interlaboratory comparison of mixtures of polystyrene with different end groups: statistical analysis of mass fractions and mass moments.

A matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) interlaboratory comparison was conducted on mixtures of synthetic polymers having the same repeat unit and closely matching molecular mass distributions but with different end groups. The interlaboratory comparison was designed to see how well the results from a group of experienced laboratories would agree on the mass fraction, and molecular mass distribution, of each polymer in a series of binary mixtures. Polystyrenes of a molecular mass near 9000 u were used. Both polystyrenes were initiated with the same butyl initiator; however, one was terminated with -H (termed PSH) and the other was terminated with -CH2CH2OH (termed PSOH). End group composition of the individual polymers was checked by MALDI-TOF MS and by nuclear magnetic resonance (NMR). Five mixtures were created gravimetrically with mass ratios between 95:5 and 10:90 PSOH/PSH. Mixture compositions where measured by NMR and by Fourier transform infrared spectrometry (FT-IR). NMR and FT-IR were used to benchmark the performance of these methods in comparison to MALDI-TOF MS. Samples of these mixtures were sent to any institution requesting it. A total of 14 institutions participated. Analysis of variance was used to examine the influences of the independent parameters (participating laboratory, MALDI matrix, instrument manufacturer, TOF mass separation mode) on the measured mass fractions and molecular mass distributions for each polymer in each mixture. Two parameters, participating laboratory and instrument manufacturer, were determined to have a statistically significant influence. MALDI matrix and TOF mass separation mode (linear or reflectron) were found not to have a significant influence. Improper mass calibration, inadequate instrument optimization with respect to high signal-to-noise ratio across the entire mass range, and poor data analysis methods (e.g., baseline subtraction and peak integration) seemed to be the greatest obstacles in the correct application of MALDI-TOF MS to this problem. Each of these problems can be addressed with proper laboratory technique.

Journal Article↗