PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Software Validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

A computer-based interview system for patients with back pain. A validation study.

A microcomputer-based system has been designed to interview patients with a view to investigating and establishing common syndromes of back and leg pain. In a randomized crossover validation study, 50 consecutive outpatients were interviewed by the computer and had a conventional clerking by a doctor. The conventional clerking made minor errors in 3.75% of questions answered and major errors in 0.90%. The computer made minor errors in 6.75% of questions and major errors in 5.45%. The majority of the computer errors were due to inadequate question design. These have been corrected, and it is anticipated that the computer will now have an overall rate of 94% correct answers and be sufficiently accurate to pursue the aim of clinical syndrome identification.

Back Pain↗

The quest for bioisosteric replacements.

To help advance drug discovery projects, a new and validated search method is presented by which potential bioisosteric replacements can be retrieved from a database of more than 700,000 structural fragments. The heart of the search method is an optimized topological pharmacophore fingerprint which describes each fragment as a combination of attachment points, hydrogen bond donors and acceptors, hydrophobic centers, conjugated atoms, and non-hydrogen atoms. In the fingerprint the influence of the attachment point is enhanced by giving it extra weight relative to the other descriptors. The Euclidean distance has proven to be the optimum distance measure to compare the fingerprints in a database search. The performance of the pharmacophore fingerprint based search method has been validated using more than 2200 bioisosteric fragment pairs extracted in an unbiased procedure from the BIOSTER database. The true bioisosteric pairs have been compared with pairs of random fragments originating from the WDI database. Normalized by the standard deviation of the random pairs distance distributions, an excellent separation of true pairs from random pairs was obtained for R-group fragments (2.2 standard deviation units) as well as for linkers (2.6 units) and cores (2.6 units). The bioisoster search method has been implemented as an intranet application called IBIS and is now routinely used by Organon researchers.

Algorithms↗

Validation of radiographic simulation codes including x-ray phase effects for millimeter-size objects with micrometer structures.

The mix between x-ray phase and attenuation information needs to be understood for accurate object recovery from radiography and tomography data. We are researching and experimentally validating algorithms that simulate x-ray phase contrast to determine the required physics necessary for quantitative object recovery. The results of a study are described to determine if a multislice (beam-propagation) method is required for simulating x-ray radiographs. We conclude that the multislice method is not required for accurate simulation of greater than or equal to 8 keV x-ray radiographs of millimeter-size objects with micrometer structures.

Algorithms↗

A surface-marker imaging system to measure a moving knee's rotational axis pathway in the sagittal plane.

This paper describes a new surface-marker imaging system designed to measure the rotational axis pathway (RAP) of a moving knee in the sagittal plane. Measurement of this parameter can provide important information about a knee's slipping and rolling action that can aid clinical assessment. Seated subjects are video recorded as they actively extend their legs. A series of stills is then captured and analyzed to extract the coordinates of markers placed on the subjects upper and lower legs. These coordinates are then processed to deduce an instant center of rotation (ICR) for each still. These ICRs are then plotted to derive the joint's RAP. The system has been validated with a mechanical model and tested in a clinical study of ten patients with unilateral anterior cruciate ligament (ACL) ruptures. The study found that the system could consistently measure differences between a patients normal and injured knees. Leg extension caused the normal knees ICRs to displace anteriorly with a mean value of 17.4 mm, whereas the injured knees had a mean displacement of 7.5 mm. This loss of roll in the ACL-deficient knees is consistent with their abnormal biomechanical arrangement.

Adult↗

Computer analysis of monophasic action potentials: manual validation and clinically pertinent applications.

Monophasic action potential (MAP) recordings are increasingly being used in a variety of clinical and experimental situations but their manual measurement is cumbersome, especially when hundreds or thousands of beats must be analyzed to monitor the exact time course of action potential duration (APD) changes following heart rate alterations, during surveillance of APD alternans, or during the onset and stabilization of Class III drug effects. To facilitate this task we developed a computer program that automates programmed electrical stimulation, digitizes at 1-kHz sampling frequency MAP recordings up to 8 channels simultaneously, analyzes all APDs at repolarization levels from 10%-90% in 10% decrements (APD10-90), and automatically outputs the analyzed numerical data into spreadsheets for graphical display or statistical analysis. To validate the computer algorithm, two independent observers manually analyzed 585 concurrent MAP recordings at a paper speed of 100 mm/s. Cycle length measurements by the computer were precise to 0.4 +/- 0.5 ms as compared to the computer determined paced cycle length. Computer measurements of APD20, 50, and 90 differed from manual measurements by 2.0 +/- 8.8 ms, 0.7 +/- 7.9 ms, and 0.2 +/- 8.5 ms, respectively, for observer 1; and by 12.2 +/- 8.3 ms, 5.8 +/- 7.5 ms, and 1.4 +/- 10.1 ms, respectively, for observer 2. Inter-observer variability (IOV) was 10.3 +/- 11.1 (APD20), 5.1 +/- 9.0 ms (APD50), and 1.2 +/- 7.8 ms (APD90), which was similar to computer/observer-2 differences and significantly greater (0.001) than computer/observer-1 differences. This indicates that the computer analysis was at least as precise as manual measurements when compared to IOV, and more precise when comparing computer/observer-1 differences to IOV. While providing equal or greater precision, computer-aided analysis of 100 MAP signals took approximately 1 minute while manual analysis of the same data set took between 2.5 and 4 hours. The pacing and analysis software was subsequently applied to experiments that mimic clinically pertinent examples of MAP recordings: (1) automatic generation, analysis, and graphical display of electrical restitution curves at multiple ventricular sites simultaneously; (2) evaluation of myocardial pharmacokinetics by monitoring the progression of Class III antiarrhythmic drug effects by continuous MAP recordings, and displaying differences in drug action between multiple sites; (3) depiction of the adaptation time course of APD to abrupt changes in paced cycle length; and (4) quantitative analysis of APD alternans during myocardial ischemia. The results show that our computerized algorithm greatly facilitates the generation of cardiac electrophysiological, and clinically important, data.

Action Potentials↗

Validating a decision support system for anti-epileptic drug treatment. Part II: adjusting anti-epileptic drug treatment.

A model of expertise for monitoring antiepileptic drug treatment was implemented in a decision support system. We validated the advice of the system regarding treatment decisions at first follow-up with 265 paper cases based on patient records. The reference for comparison is based on the opinions of neurologists. We found considerable variation among the decisions of five neurologists. It could be shown that the system agreed with (groups of) neurologists at least as often as individual neurologists did. The correctness of the system was consistently higher than that of each of the neurologists, when the majority decision of the remaining neurologists constituted the standard.

Anticonvulsants↗

Comparison of automatic quantification software for the measurement of ventricular volume and ejection fraction in gated myocardial perfusion SPECT.

The aim of this study was to compare the performance of three different software packages for the calculation of ejection fraction (EF) and end diastolic volume (EDV) from gated myocardial single photon emission computed tomography studies. Two hundred patients undergoing gated stress myocardial perfusion scans were analysed retrospectively. Patients were grouped as follows: small heart (n=31), normal perfusion scan (n=71), and scan with perfusion defects (n=98). EF and EDV were calculated for each using QGS (Cedars Sinai, Los Angeles, CA), 4D-MSPECT (University of Michigan, Ann Arbor, MI), and ECT (Emory University, Atlanta, GA). Bland-Altman plots, repeated measures ANOVA, and linear regression analysis were used to compare methods. Correlation coefficients between the methods for both EF and EDV were high, greater than 0.9. However, Bland-Altman plots revealed a large standard deviation of the difference between methods, preventing the confident estimate of the value of one method from an observation of another. Despite good correlation, the variance between methods was high. These algorithms behave differently, produce widely variable results from one another, and should not be used interchangeably. It may prove prudent for laboratories to independently validate the software algorithm that is chosen against a 'gold standard' using their own population.

Adolescent↗

On the assets of CAD planning for craniosynostosis surgery.

SkullWiz is a computer-aided design program that transforms computer tomographic data of the neurocranium into a mathematical model that can be interactively manipulated to plan craniosynostosis surgery. Proper planning of this type of surgery involves reference to the underlying viscerocranium and to normal neurocranial dimensions, simulation of all basic surgical actions (closed and open osteotomy, translation, rotation, bending, removal, burring), and reference to the mechanical properties of calvarial bone at a given age. With SkullWiz, infinite trials are possible to develop a surgical plan that combines minimal action with maximum morphologic result. In contrast, physical models, e.g., foam milled or stereolitographic, provide just a single (or double, after gluing) opportunity to visualize three-dimensional morphology and simulate a treatment plan, without reference support. Validation of SkullWiz is difficult due to parameter variability. Its assets are therefore graphically exemplified in two common types of nonsyndromatic single-suture craniosynostosis-trigonocephaly and anterior plagiocephaly. SkullWiz is one of the most accurate planning tools currently available for craniosynostosis surgery. Accurate transfer of the planning by aluminium templates results in efficient and precise surgery by avoiding per-operative "chipping and fitting."

Age Factors↗

Detecting errors in a scoring program: a method of double diagnosis using a computer-generated sample.

This paper discusses a new method for locating errors in diagnostic computer scoring programs for structured clinical interviews. It was proposed as a test of the accuracy of the scoring program for the Composite International Diagnostic Interview, version 1.1. The proposal was to create an independent scoring program in a different computer language but serving the same criteria. Both programs were then applied to the same large set of valid (i.e., logically consistent) computer-generated test cases, and differences in diagnostic assignments reviewed. The method described can identify the program steps that account for the sources of the errors. Corrections can be made and the programs run again on new sets of test cases until discrepancy-free results are achieved. While this method cannot discover errors that are repeated in the two programs, it does discover more of the errors in a scoring program than we have previously been able to identify. This technique provides a systematic and rigorous approach to assuring the accuracy of scoring programs based on established algorithms.

Algorithms↗

Validation of a Monte Carlo simulation of the Philips Allegro/GEMINI PET systems using GATE.

A newly developed simulation toolkit, GATE (Geant4 Application for Tomographic Emission), was used to develop a Monte Carlo simulation of a fully three-dimensional (3D) clinical PET scanner. The Philips Allegro/GEMINI PET systems were simulated in order to (a) allow a detailed study of the parameters affecting the system's performance under various imaging conditions, (b) study the optimization and quantitative accuracy of emission acquisition protocols for dynamic and static imaging, and (c) further validate the potential of GATE for the simulation of clinical PET systems. A model of the detection system and its geometry was developed. The accuracy of the developed detection model was tested through the comparison of simulated and measured results obtained with the Allegro/GEMINI systems for a number of NEMA NU2-2001 performance protocols including spatial resolution, sensitivity and scatter fraction. In addition, an approximate model of the system's dead time at the level of detected single events and coincidences was developed in an attempt to simulate the count rate related performance characteristics of the scanner. The developed dead-time model was assessed under different imaging conditions using the count rate loss and noise equivalent count rates performance protocols of standard and modified NEMA NU2-2001 (whole body imaging conditions) and NEMA NU2-1994 (brain imaging conditions) comparing simulated with experimental measurements obtained with the Allegro/GEMINI PET systems. Finally, a reconstructed image quality protocol was used to assess the overall performance of the developed model. An agreement of <3% was obtained in scatter fraction, with a difference between 4% and 10% in the true and random coincidence count rates respectively, throughout a range of activity concentrations and under various imaging conditions, resulting in <8% differences between simulated and measured noise equivalent count rates performance. Finally, the image quality validation study revealed a good agreement in signal-to-noise ratio and contrast recovery coefficients for a number of different volume spheres and two different (clinical level based) tumour-to-background ratios. In conclusion, these results support the accurate modelling of the Philips Allegro/GEMINI PET systems using GATE in combination with a dead-time model for the signal flow description, which leads to an agreement of <10% in coincidence count rates under different imaging conditions and clinically relevant activity concentration levels.

Computer Simulation↗

Utilization of two sample t-test statistics from redundant probe sets to evaluate different probe set algorithms in GeneChip studies.

BACKGROUND: The choice of probe set algorithms for expression summary in a GeneChip study has a great impact on subsequent gene expression data analysis. Spiked-in cRNAs with known concentration are often used to assess the relative performance of probe set algorithms. Given the fact that the spiked-in cRNAs do not represent endogenously expressed genes in experiments, it becomes increasingly important to have methods to study whether a particular probe set algorithm is more appropriate for a specific dataset, without using such external reference data. RESULTS: We propose the use of the probe set redundancy feature for evaluating the performance of probe set algorithms, and have presented three approaches for analyzing data variance and result bias using two sample t-test statistics from redundant probe sets. These approaches are as follows: 1) analyzing redundant probe set variance based on t-statistic rank order, 2) computing correlation of t-statistics between redundant probe sets, and 3) analyzing the co-occurrence of replicate redundant probe sets representing differentially expressed genes. We applied these approaches to expression summary data generated from three datasets utilizing individual probe set algorithms of MAS5.0, dChip, or RMA. We also utilized combinations of options from the three probe set algorithms. We found that results from the three approaches were similar within each individual expression summary dataset, and were also in good agreement with previously reported findings by others. We also demonstrate the validity of our findings by independent experimental methods. CONCLUSION: All three proposed approaches allowed us to assess the performance of probe set algorithms using the probe set redundancy feature. The analyses of redundant probe set variance based on t-statistic rank order and correlation of t-statistics between redundant probe sets provide useful tools for data variance analysis, and the co-occurrence of replicate redundant probe sets representing differentially expressed genes allows estimation of result bias. The results also suggest that individual probe set algorithms have dataset-specific performance.

Algorithms↗

Boosted leave-many-out cross-validation: the effect of training and test set diversity on PLS statistics.

It is becoming increasingly common in quantitative structure/activity relationship (QSAR) analyses to use external test sets to evaluate the likely stability and predictivity of the models obtained. In some cases, such as those involving variable selection, an internal test set--i.e., a cross-validation set--is also used. Care is sometimes taken to ensure that the subsets used exhibit response and/or property distributions similar to those of the data set as a whole, but more often the individual observations are simply assigned 'at random.' In the special case of MLR without variable selection, it can be analytically demonstrated that this strategy is inferior to others. Most particularly, D-optimal design performs better if the form of the regression equation is known and the variables involved are well behaved. This report introduces an alternative, non-parametric approach termed 'boosted leave-many-out' (boosted LMO) cross-validation. In this method, relatively small training sets are chosen by applying optimizable k-dissimilarity selection (OptiSim) using a small subsample size (k = 4, in this case), with the unselected observations being reserved as a test set for the corresponding reduced model. Predictive errors for the full model are then estimated by aggregating results over several such analyses. The countervailing effects of training and test set size, diversity, and representativeness on PLS model statistics are described for CoMFA analysis of a large data set of COX2 inhibitors.

Algorithms↗

Machine learning for an expert system to predict preterm birth risk.

OBJECTIVE: Develop a prototype expert system for preterm birth risk assessment of pregnant women. Normal gestation involves a term of 40 weeks, but because 8-12% of the newborns in the United States are delivered prior to 37 weeks' gestation, problems associated with prematurity continue to plague individuals, families, and the health care system. DESIGN: A knowledge-base development methodology used machine learning, statistical analysis, and validation techniques to analyze three large datasets (18,890 subjects and 214 variables). The dependent (i.e., decision) variable studied was weeks of gestation at delivery, with dichotomous coding of preterm delivery (prior to 37 weeks) and full-term delivery (37+ weeks). RESULTS: Machine learning with a program named Learning from Examples using Rough Sets (LERS) induced 520 usable rules that were entered into a prototype expert system. The prototype expert system was 53-88% accurate in predicting preterm delivery for 9,419 patients. CONCLUSION: The prototype expert system was more accurate than traditional manual techniques in predicting preterm birth.

Adult↗

Preparation of name and address data for record linkage using hidden Markov models.

BACKGROUND: Record linkage refers to the process of joining records that relate to the same entity or event in one or more data collections. In the absence of a shared, unique key, record linkage involves the comparison of ensembles of partially-identifying, non-unique data items between pairs of records. Data items with variable formats, such as names and addresses, need to be transformed and normalised in order to validly carry out these comparisons. Traditionally, deterministic rule-based data processing systems have been used to carry out this pre-processing, which is commonly referred to as "standardisation". This paper describes an alternative approach to standardisation, using a combination of lexicon-based tokenisation and probabilistic hidden Markov models (HMMs). METHODS: HMMs were trained to standardise typical Australian name and address data drawn from a range of health data collections. The accuracy of the results was compared to that produced by rule-based systems. RESULTS: Training of HMMs was found to be quick and did not require any specialised skills. For addresses, HMMs produced equal or better standardisation accuracy than a widely-used rule-based system. However, accuracy was worse when used with simpler name data. Possible reasons for this poorer performance are discussed. CONCLUSION: Lexicon-based tokenisation and HMMs provide a viable and effort-effective alternative to rule-based systems for pre-processing more complex variably formatted data such as addresses. Further work is required to improve the performance of this approach with simpler data such as names. Software which implements the methods described in this paper is freely available under an open source license for other researchers to use and improve.

Data Collection↗

Low-energy imaging with high-energy bremsstrahlung beams: analysis and scatter reduction.

The contrast and zero spatial frequency signal-to-noise ratio produced by a method for radiation therapy portal imaging known as low-energy imaging with high-energy bremsstrahlung beams have been mathematically analyzed. The analysis makes extensive use of Monte Carlo techniques and incorporates the detector, the spectrum, phantom, and geometry. The analysis is validated through comparison with measured data including subject contrast measurements and the attenuation of the beam with lead. Scatter reduction is found to be potentially the most effective method to improve contrast and SNR for a film based system. A large fraction of the scatter detected is of a much higher energy than that found in diagnostic radiology. Hence, traditional antiscatter grids, such as those used in diagnostic radiology, are ineffective. The analysis and theory from the literature are applied to design a new grid which is more appropriate for this application. The grid produces a modest improvement according to a contrast-detail study.

Humans↗

A Bayesian method for identifying missing enzymes in predicted metabolic pathway databases.

BACKGROUND: The PathoLogic program constructs Pathway/Genome databases by using a genome's annotation to predict the set of metabolic pathways present in an organism. PathoLogic determines the set of reactions composing those pathways from the enzymes annotated in the organism's genome. Most annotation efforts fail to assign function to 40-60% of sequences. In addition, large numbers of sequences may have non-specific annotations (e.g., thiolase family protein). Pathway holes occur when a genome appears to lack the enzymes needed to catalyze reactions in a pathway. If a protein has not been assigned a specific function during the annotation process, any reaction catalyzed by that protein will appear as a missing enzyme or pathway hole in a Pathway/Genome database. RESULTS: We have developed a method that efficiently combines homology and pathway-based evidence to identify candidates for filling pathway holes in Pathway/Genome databases. Our program not only identifies potential candidate sequences for pathway holes, but combines data from multiple, heterogeneous sources to assess the likelihood that a candidate has the required function. Our algorithm emulates the manual sequence annotation process, considering not only evidence from homology searches, but also considering evidence from genomic context (i.e., is the gene part of an operon?) and functional context (e.g., are there functionally-related genes nearby in the genome?) to determine the posterior belief that a candidate has the required function. The method can be applied across an entire metabolic pathway network and is generally applicable to any pathway database. The program uses a set of sequences encoding the required activity in other genomes to identify candidate proteins in the genome of interest, and then evaluates each candidate by using a simple Bayes classifier to determine the probability that the candidate has the desired function. We achieved 71% precision at a probability threshold of 0.9 during cross-validation using known reactions in computationally-predicted pathway databases. After applying our method to 513 pathway holes in 333 pathways from three Pathway/Genome databases, we increased the number of complete pathways by 42%. We made putative assignments to 46% of the holes, including annotation of 17 sequences of previously unknown function. CONCLUSIONS: Our pathway hole filler can be used not only to increase the utility of Pathway/Genome databases to both experimental and computational researchers, but also to improve predictions of protein function.

Amino Acid Oxidoreductases↗

Validating clustering for gene expression data.

MOTIVATION: Many clustering algorithms have been proposed for the analysis of gene expression data, but little guidance is available to help choose among them. We provide a systematic framework for assessing the results of clustering algorithms. Clustering algorithms attempt to partition the genes into groups exhibiting similar patterns of variation in expression level. Our methodology is to apply a clustering algorithm to the data from all but one experimental condition. The remaining condition is used to assess the predictive power of the resulting clusters-meaningful clusters should exhibit less variation in the remaining condition than clusters formed by chance. RESULTS: We successfully applied our methodology to compare six clustering algorithms on four gene expression data sets. We found our quantitative measures of cluster quality to be positively correlated with external standards of cluster quality.

Algorithms↗

Validation, clinical trial, and evaluation of a radiology expert system.

The PHOENIX Radiology Consultant is a rule-based expert system which assists physicians in planning radiological work-up strategies. This article describes the methods used to create and validate the system's knowledge base. The feasibility and acceptability of PHOENIX were tested for two years in a clinical trial. During this period, the system was used 1,421 times, an average of 13.7 times per week, primarily by medical students and nonradiologist physicians. Much of the system's use occurred at night and on weekends, when the radiology department was not fully staffed. Several physicians were enlisted to further evaluate the utility of the system. The results of their evaluation indicate that an expert system that helps physicians select diagnostic-imaging studies can serve as a useful and informative component of a radiology information system, and is particularly useful for medical students and physicians in training.

Algorithms↗