PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Performance benchmarking”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Using risk assessment to evaluate adverse selection under capitated contracts.

Most provider organizations rely on health plan or market-based information about capitation rates, per member per month costs, and utilization trends to benchmark their performance. However, these statistics can be misleading because of differences in enrollee mix and contracting terms across provider organizations. This article describes the limitations of health plan contract provisions in protecting against adverse selection. It describes various actuarial and statistical data sources for evaluation of adverse selection. The article then presents various approaches to risk adjustment on population basis and their use in quantifying adverse selection for health plan contract negotiations.

Actuarial Analysis↗

Scalable medium-density genotyping platforms for cultivar identification, pedigree authentication, marker-assisted and genomic selection, and other applications in strawberry.

A broad spectrum of high-density genotyping approaches, including single-nucleotide polymorphism (SNP) arrays, genotyping-by-sequencing, and whole-genome reduced-representation sequencing, have been shown to perform well in strawberry (Fragaria × ananassa), despite the inherent complexity of the octoploid genome. While these approaches are effective, their routine deployment in breeding programs can be constrained by cost, computational requirements, and workflow complexity. In parallel, many breeding programs continue to rely on locus-specific assays for marker-assisted selection, resulting in fragmented and inefficient genotyping strategies. Here, we describe medium-density amplicon-based genotyping platforms for strawberry designed to provide cost-effective, turnkey solutions that integrate markers used for marker-assisted selection with genome-wide markers suitable for genomic prediction in a single laboratory assay. These platforms were developed by targeting 1,650 or 4,811 target SNPs via amplicon sequencing, and are interoperable with existing high-density genotyping resources, including a widely used 50K SNP array, thereby facilitating data integration across platforms. We benchmarked their performance relative to the 50K SNP array across breeding-relevant applications, including identity and purity testing, pedigree authentication, marker-assisted selection, and genomic selection, and further evaluated the feasibility of genotype imputation to enhance genome-wide information content. Across analyses, the 1,650- and 4,811-amplicon platforms produced results comparable to higher-density platforms while substantially reducing genotyping cost and analytical overhead. This work demonstrates that targeted amplicon-based genotyping can support efficient, scalable, and integrated genome-informed breeding, enabling the routine application of both marker-assisted and genomic selection within strawberry breeding workflows. Open-source R workflows are provided to support streamlined analyses in breeding contexts.

Fragaria↗

Validating the performance of the mammary sentinel lymph node team.

BACKGROUND AND OBJECTIVES: The mammary sentinel lymph node (SLN) procedure has the potential to improve the accuracy and lower the morbidity of axillary staging in breast cancer patients, but results are closely linked to experience and can vary widely between institutions. Standardized performance measures need to be established in order to optimize the transition to SLN biopsy only. METHODS: Performance data were prospectively collected for the first 156 mammary SLN procedures performed by three surgeons in our institution. RESULTS: Seventy-five cases were required to achieve an SLN visualization rate of > 80% on preoperative lymphoscintigraphy. The SLN visualization rate was 90% for the last 52 cases. Two surgeons required 25 cases before consistently achieving a > or = 90% SLN identification rate in the operating room and one required 15 cases. The metastasis detection rate increased from 22% for the first 52 cases to 31% for the last 52 cases. The false negative rate for the procedure was 5%. CONCLUSIONS: The following performance criteria and benchmarks are suggested for validating the performance of the SLN team: (1) SLN visualization rate on preoperative lymphoscintigraphy > or = 80%, (2) SLN identification rate in the operating room > or = 90%, (3) False negative rate for the procedure 5%. Thirty procedures per surgeon were sufficient to achieve these benchmarks in our group.

Axilla↗

Automated Classification of Lymphoma Subtypes From Histopathological Images Using a U-Net Deep Learning Model: Comparative Evaluation Study.

BACKGROUND: Accurate classification and grading of lymphoma subtypes are essential for treatment planning. Traditional diagnostic methods face challenges of subjectivity and inefficiency, highlighting the need for automated solutions based on deep learning techniques. OBJECTIVE: This study aimed to investigate the application of deep learning technology, specifically the U-Net model, in classifying and grading lymphoma subtypes to enhance diagnostic precision and efficiency. METHODS: In this study, the U-Net model was used as the primary tool for image segmentation integrated with attention mechanisms and residual networks for feature extraction and classification. A total of 620 high-quality histopathological images representing 3 major lymphoma subtypes were collected from The Cancer Genome Atlas and the Cancer Imaging Archive. All images underwent standardized preprocessing, including Gaussian filtering for noise reduction, histogram equalization, and normalization. Data augmentation techniques such as rotation, flipping, and scaling were applied to improve the model's generalization capability. The dataset was divided into training (70%), validation (15%), and test (15%) subsets. Five-fold cross-validation was used to assess model robustness. Performance was benchmarked against mainstream convolutional neural network architectures, including fully convolutional network, SegNet, and DeepLabv3+. RESULTS: The U-Net model achieved high segmentation accuracy, effectively delineating lesion regions and improving the quality of input for classification and grading. The incorporation of attention mechanisms further improved the model's ability to extract key features, whereas the residual structure of the residual network enhanced classification accuracy for complex images. In the test set (N=1250), the proposed fusion model achieved an accuracy of 92% (1150/1250), a sensitivity of 91.04% (1138/1250), a specificity of 89.04% (1113/1250), and an F1-score of 90% (1125/1250) for the classification of the 3 lymphoma subtypes, with an area under the receiver operating characteristic curve of 0.95 (95% CI 0.93-0.97). The high sensitivity and specificity of the model indicate strong clinical applicability, particularly as an assistive diagnostic tool. CONCLUSIONS: Deep learning techniques based on the U-Net architecture offer considerable advantages in the automated classification and grading of lymphoma subtypes. The proposed model significantly improved diagnostic accuracy and accelerated pathological evaluation, providing efficient and precise support for clinical decision-making. Future work may focus on enhancing model robustness through integration with advanced algorithms and validating performance across multicenter clinical datasets. The model also holds promise for deployment in digital pathology platforms and artificial intelligence-assisted diagnostic workflows, improving screening efficiency and promoting consistency in pathological classification.

Humans↗

Evaluation and development of potentially better practices to prevent chronic lung disease and reduce lung injury in neonates.

OBJECTIVE: Despite increased knowledge and improving technology, chronic lung disease (CLD) rates in extremely low birth weight infants have remained constant for 20 years. One reason for this is an ineffective translation of research-proven improvements into practice. The Neonatal Intensive Care Quality Improvement Collaborative Year 2000 (NIC/Q 2000) was created to provide participating nurseries the tools necessary to effect change. The objective of this study was to develop and implement a process that uses quality improvement techniques to collaboratively improve CLD rates. METHODS: Nine member hospitals of the NIC/Q 2000 collaborative formed a focus group aiming to decrease CLD rates. The focus group established goals and outcome measures, created a list of potentially better practices (PBPs) based on available literature, benchmarked and performed site visits, encouraged individual site implementation of PBPs, developed a database, and measured outcomes. RESULTS: The goal "decrease CLD rates in extremely low birth weight infants" was established. Nine PBPs were identified, and 57 PBPs were implemented by the 9 participating sites. Twelve site visits were conducted, and a 435-patient database of infants with a mean birth weight of 789 g was established. CONCLUSIONS: Collaborative use of quality improvement techniques resulted in creation of a logical, efficient, and effective process to improve CLD rates. Group creation of PBPs, based on literature review and reinforced with site visits, internal data analysis, and improved individual site outcomes, resulted in accelerated and effective change, unlikely to occur if attempted outside of the collaborative.

Benchmarking↗

In vivo dehydration of silicone hydrogel contact lenses.

PURPOSE: To benchmark the performance of new-generation silicone hydrogel contact lenses in terms of their in vivo hydration characteristics and to highlight the possible clinical ramifications of any changes observed. METHODS: Thirteen subjects (four men and nine women with a mean age of 24.8 +/- 5.0 years) wore a silicone hydrogel lens (PureVision, Bausch & Lomb, Rochester, NY) in one eye and a conventional hydrogel lens (ACUVUE 2, Johnson & Johnson Vision Care, Jacksonville, FL) in the other eye for 4 weeks on an extended-wear basis. A gravimetric method was used to determine lens water content and dehydration during the intended lifespan of the lenses. RESULTS: For the PureVision lens, the water content was 38.3% +/- 0.9%, 35.2% +/- 1.1%, and 35.3% +/- 1.7% at baseline and after 2 and 4 weeks of wear, respectively (F=28.4, P<0.0001). For the ACUVUE 2 lens, the water content was 58.1% +/- 0.6% and 52.1% +/- 1.3% at baseline and after 2 weeks of wear, respectively. Thus, after 2 weeks of wear, absolute dehydration was 2.8% +/- 1.8% and 6.0% +/- 1.3% for the PureVision and ACUVUE 2 lenses, respectively (t=6.8, P<0.0001). The mass of deposition was calculated to be 568 +/- 457 microg for the PureVision lens and 1,660 +/- 499 microg for the ACUVUE 2 lens (t=5.1, P=0.0003). CONCLUSIONS: The ACUVUE 2 lens underwent a greater degree of lens dehydration, causing a reduction in oxygen permeability (3.6 barrer), and deposition after 2 weeks of extended wear. The loss of water from the PureVision lens was paradoxically associated with a 6.0-barrer increase in oxygen permeability after 4 weeks of extended wear.

Adult↗

Top-performing physician groups focus on innovative cost control.

DATA BENCHMARKS: Top-performing physician group practices reveal why they are head and shoulders above their peers. A new study released by the Medical Group Management Association offers benchmark data that shows how leading group practices in various specialties are faring compared with their colleagues nationwide.

Benchmarking↗

A branch and bound algorithm for protein structure refinement from sparse NMR data sets.

We describe new methods for predicting protein tertiary structures to low resolution given the specification of secondary structure and a limited set of long-range NMR distance constraints. The NMR data sets are derived from a realistic protocol involving completely deuterated 15N and 13C-labeled samples. A global optimization method, based upon a modification of the alphaBB (branch and bound) algorithm of Floudas and co-workers, is employed to minimize an objective function combining the NMR distance restraints with a residue-based protein folding potential containing hydrophobicity, excluded volume, and van der Waals interactions. To assess the efficacy of the new methodology, results are compared with benchmark calculations performed via the X-PLOR program of Brünger and co-workers using standard distance geometry/molecular dynamics (DGMD) calculations. Seven mixed alpha/beta proteins are examined, up to a size of 183 residues, which our methods are able to treat with a relatively modest computational effort, considering the size of the conformational space. In all cases, our new approach provides substantial improvement in root-mean-square deviation from the native structure over the DGMD results; in many cases, the DGMD results are qualitatively in error, whereas the new method uniformly produces high quality low-resolution structures. The DGMD structures, for example, are systematically non-compact, which probably results from the lack of a hydrophobic term in the X-PLOR energy function. These results are highly encouraging as to the possibility of developing computational/NMR protocols for accelerating structure determination in larger proteins, where data sets are often underconstrained.

Algorithms↗

Mining gene expression data using a novel approach based on hidden Markov models.

In this work we have developed a new framework for microarray gene expression data analysis. This framework is based on hidden Markov models. We have benchmarked the performance of this probability model-based clustering algorithm on several gene expression datasets for which external evaluation criteria were available. The results showed that this approach could produce clusters of quality comparable to two prevalent clustering algorithms, but with the major advantage of determining the number of clusters. We have also applied this algorithm to analyze published data of yeast cell cycle gene expression and found it able to successfully dig out biologically meaningful gene groups. In addition, this algorithm can also find correlation between different functional groups and distinguish between function genes and regulation genes, which is helpful to construct a network describing particular biological associations. Currently, this method is limited to time series data. Supplementary materials are available at http://www.bioinfo.tsinghua.edu.cn/~rich/hmmgep_supp/.

Algorithms↗

A novel learning algorithm which improves the partial fault tolerance of multilayer neural networks.

The paper deals with the problem of fault tolerance in a multilayer perceptron network. Although it already possesses a reasonable fault tolerance capability, it may be insufficient in particularly critical applications. Studies carried out by the authors have shown that the traditional backpropagation learning algorithm may entail the presence of a certain number of weights with a much higher absolute value than the others. Further studies have shown that faults in these weights is the main cause of deterioration in the performance of the neural network. In other words, the main cause of incorrect network functioning on the occurrence of a fault is the non-uniform distribution of absolute values of weights in each layer. The paper proposes a learning algorithm which updates the weights, distributing their absolute values as uniformly as possible in each layer. Tests performed on benchmark test sets have shown the considerable increase in fault tolerance obtainable with the proposed approach as compared with the traditional backpropagation algorithm and with some of the most efficient fault tolerance approaches to be found in literature.

Journal Article↗

Towards efficient perturbation for the noncoding genome.

Deciphering the functionality of the noncoding genome, which includes important cis-regulatory elements (CREs) and transcribed noncoding RNA genes, remains technically challenging. Here, using massively parallel genetic screening, we systematically benchmark the performance of five representative loss-of-function perturbation tools, including single-guide RNA (gRNA) mediated SpCas9 cleavage or CRISPR interference, and paired gRNA (pgRNA) involved dual-SpCas9, Big Papi (paired SpCas9 and SaCas9) or dual-enAsCas12a fragment deletion methods, in decoding the roles of the noncoding genome. For targeting CREs such as enhancers, dual-SpCas9 outperforms other methods with superior efficiency in destroying functional genomic regions. For perturbing noncoding RNA genes, in addition to dual-SpCas9, other RNA-targeting methods such as RNA interference are recommended to discriminate transcript-dependent or -independent roles. A deep learning model, DeepDC, with an associated web server, is built to facilitate optimal dual-SpCas9 pgRNA design for efficiently deleting a genomic fragment. Together, our work provides practical guidance on selecting appropriate loss-of-function tools to resolve the functional complexity of the noncoding genome.

CRISPR-Cas Systems↗

Development of a model for prediction of survival in pediatric trauma patients: comparison of artificial neural networks and logistic regression.

BACKGROUND/PURPOSE: There is a paucity of outcome prediction models for injured children. Using the National Pediatric Trauma Registry (NPTR), the authors developed an artificial neural network (ANN) to predict pediatric trauma death and compared it with logistic regression (LR). METHODS: Patients in the NPTR from 1996 through 1999 were included. Models were generated using LR and ANN. A data search engine was used to generate the ANN with the best fit for the data. Input variables included anatomic and physiologic characteristics. There was a single output variable: probability of death. Assessment of the models was for both discrimination (ROC area under the curve) and calibration (Lemeshow-Hosmer C-Statistic). RESULTS: There were 35,385 patients. The average age was 8.1 +/- 5.1 years, and there were 1,047 deaths (3.0%). Both modeling systems gave excellent discrimination (ROC A(z): LR = 0.964, ANN = 0.961). However, LR had only fair calibration, whereas the ANN model had excellent calibration (L/H C stat: LR = 36, ANN = 10.5). CONCLUSIONS: The authors were able to develop an ANN model for the prediction of pediatric trauma death, which yielded excellent discrimination and calibration exceeding that of logistic regression. This model can be used by trauma centers to benchmark their performance in treating the pediatric trauma population.

Calibration↗

Model-based clustering and data transformations for gene expression data.

MOTIVATION: Clustering is a useful exploratory technique for the analysis of gene expression data. Many different heuristic clustering algorithms have been proposed in this context. Clustering algorithms based on probability models offer a principled alternative to heuristic algorithms. In particular, model-based clustering assumes that the data is generated by a finite mixture of underlying probability distributions such as multivariate normal distributions. The issues of selecting a 'good' clustering method and determining the 'correct' number of clusters are reduced to model selection problems in the probability framework. Gaussian mixture models have been shown to be a powerful tool for clustering in many applications. RESULTS: We benchmarked the performance of model-based clustering on several synthetic and real gene expression data sets for which external evaluation criteria were available. The model-based approach has superior performance on our synthetic data sets, consistently selecting the correct model and the number of clusters. On real expression data, the model-based approach produced clusters of quality comparable to a leading heuristic clustering algorithm, but with the key advantage of suggesting the number of clusters and an appropriate model. We also explored the validity of the Gaussian mixture assumption on different transformations of real data. We also assessed the degree to which these real gene expression data sets fit multivariate Gaussian distributions both before and after subjecting them to commonly used data transformations. Suitably chosen transformations seem to result in reasonable fits. AVAILABILITY: MCLUST is available at http://www.stat.washington.edu/fraley/mclust. The software for the diagonal model is under development. CONTACT: kayee@cs.washington.edu. SUPPLEMENTARY INFORMATION: http://www.cs.washington.edu/homes/kayee/model.

Algorithms↗

Comprehensive Evaluation and Explainable Interpretation of Peptide-HLA Binding Prediction Tools.

Accurate prediction of peptide binding to human leukocyte antigen class I (HLA-I) molecules is critical for advancing immunological research, particularly in vaccine design and immunotherapy. However, limitations in model performance, interpretability, and dataset quality impede the widespread adoption of existing predictive tools. Here, we present a comprehensive evaluation of 17 HLA-I peptide binding prediction models, utilizing a meticulously curated dataset comprising over 290,000 peptides spanning 44 HLA-I alleles. We assessed model accuracy, robustness, and interpretability, employing explainability techniques such as SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) to elucidate underlying prediction mechanisms. Our results reveal substantial performance disparities, with self-attention-based models, including STMHCpan and BigMHC, exhibiting superior accuracy. Notably, the capsule network model CapsNet-MHC_AN demonstrated robust performance. Models trained on eluted ligand datasets outperformed those relying on binding affinity data, underscoring the critical role of high-quality training data. Ensemble and multi-algorithm approaches further improved prediction reliability. These findings highlight the need for ongoing innovation in model architecture, integration of diverse and high-quality datasets, and incorporation of structural predictors to develop more accurate, interpretable, and clinically applicable HLA-I peptide binding prediction tools.

HLA-I binding↗

Genome-resolved assessment of archaeal diversity in full-scale anaerobic digesters reveals variability in mcrA primer coverage.

AIMS: Methanogenic archaea are key players in anaerobic digestion, driving methane production in biogas reactors. This study aimed to assess the diversity of methanogenic archaea in full-scale anaerobic digesters using genome-resolved metagenomics and to systematically evaluate the taxonomic coverage of commonly used mcrA-targeted qPCR primer sets against this genomic framework. METHODS AND RESULTS: Methanogenic diversity was assessed using 113 dereplicated archaeal metagenome-assembled genomes (MAGs) recovered from 109 full-scale anaerobic digesters treating diverse substrates. Genome-resolved analyses revealed a diverse archaeal community spanning multiple phyla, dominated by Halobacteriota and Methanobacteriota, with additional representatives from Methanobacteriota_B, Thermoplasmatota, and Thermoproteota. The presence of the mcrA gene was identified in a subset 55 MAGs, which were subsequently used as the genomic framework to evaluate six commonly used mcrA qPCR primer sets in silico. This subset clustered into nine phylogenetic groups and formed the basis for the primer coverage analysis. The evaluation revealed marked differences in taxonomic coverage among primer sets. Most primers preferentially detected Methanobacteriales and Methanosarcinales, while underrepresenting or excluding other methanogenic lineages, including H&#x2082;-dependent methylotrophic Methanomassiliicoccaceae. CONCLUSIONS: Commonly used mcrA primer sets differ substantially in their ability to capture methanogenic diversity, with some showing broad representation of reactor-associated methanogens and others exhibiting strong lineage-specific biases. Genome-resolved metagenomics provides an effective framework for benchmarking primer performance and supports the selection and improvement of molecular tools for more accurate monitoring of anaerobic digestion systems.

Archaea↗

Governing the borderlands: decoding the power of aid.

This article examines aid practice, that is, the public-private contractual networks that link donor governments, UN agencies, military establishments, NGOs, private companies and others, as a relation of global liberal governance. In order to fulfil this function, such networks embody what could be called the 'securitisation' of international assistance. Based upon ideas of human security and ameliorating the effects of poverty and vulnerability reduction, aid is now seen as playing a direct security role. Rather than being concerned with relations between states, the primary aim of this security paradigm is to modulate and change the behaviour of populations within them. In doing so, it is able to exploit the opportunities afforded by privatisation. At the same time, however, aid as security is confronted by its own particular problem of 'governing at a distance'; how can calculations made by leading states be transformed into actions at the global edge when a multitude of private and non-government implementors now intervene? The article concludes by examining the contribution of risk analysis to solving this problem and, especially, the development of new contractual regimes based around technical standardisation, benchmarking and performance auditing. Through such technologies, metropolitan states are learning how to manage the public-private networks of aid practice and, as a result, to govern the borderlands in new ways.

Humans↗

Tuning diversity in bagged ensembles.

In this paper, we investigate how the level of diversity amongst individual neural networks in a bagged ensemble can significantly influence overall ensemble generalization performance. We propose a new technique that tunes this diversity so that ensemble generalization performance is optimized and evaluate its performance on benchmark regression data-sets.

Algorithms↗

Computed tomography-guided precision biopsy combined with metagenomic next-generation sequencing for etiological diagnosis in patients with blood culture-negative systemic infections.

ObjectiveTo evaluate the diagnostic efficacy of computed tomography-guided percutaneous biopsy combined with metagenomic next-generation sequencing in patients with blood culture-negative systemic infections and to assess the clinical impact of using this combined strategy for etiological confirmation and guidance of targeted antimicrobial therapy.MethodsThis single-center retrospective observational cohort study enrolled 78 patients who met the Sepsis-3 consensus criteria for suspected systemic infection and had negative conventional microbiological work-ups (at least two sets of blood cultures) between April 2022 and March 2025. All patients underwent computed tomography-guided biopsy of radiologically identified infectious foci, with specimens processed concurrently for conventional culture and metagenomic next-generation sequencing. Diagnostic performance was benchmarked against the final comprehensive clinical diagnosis, and the influence of metagenomic next-generation sequencing findings on antimicrobial therapy modification was analyzed. Sample size calculation, based on a prior study estimating an metagenomic next-generation sequencing detection rate of 85% (&#x3b1;&#x2009;=&#x2009;0.05, &#x3b2;&#x2009;=&#x2009;0.2), indicated a minimum of 68 cases; accordingly, 78 patients were enrolled.ResultsComputed tomography-guided biopsy was technically successful in all 78 patients (100%). The pathogen detection rate of metagenomic next-generation sequencing (91.0%, 71/78) was significantly higher than that of conventional culture (55.1%, 43/78; p&#x2009;<&#x2009;0.001). Using the final clinical diagnosis as the reference standard, metagenomic next-generation sequencing achieved a sensitivity of 94.7% (95% confidence interval: 86.9-98.5), specificity of 100.0% (95% confidence interval: 29.2-100.0), positive predictive value of 100.0% (95% confidence interval: 94.9-100.0), and negative predictive value of 42.9% (95% confidence interval: 9.9-81.6). Among the 35 culture-negative specimens, metagenomic next-generation sequencing established a definitive microbiological diagnosis in 28 cases (80.0%) and detected polymicrobial infections in 11 cases (14.1% of the cohort). Antimicrobial therapy was rationally adjusted based on metagenomic next-generation sequencing results in 69.2% (54/78) of the patients.ConclusionsThe integration of computed tomography-guided precision biopsy with metagenomic next-generation sequencing offers a highly effective diagnostic approach for blood culture-negative systemic infections. This synergistic strategy improves etiological diagnosis by providing high-yield target specimens that enable comprehensive, unbiased pathogen screening, facilitates differentiation between infectious and non-infectious etiologies, and supplies critical evidence for guiding precision antimicrobial therapy. These findings highlight the growing role of interventional radiology in the contemporary framework of precision infectious disease management.

Humans↗