PubMed HealthSearch

SEARCH · PubMed Health

Results for “Software Validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Methodological issues in validating decision-support systems for insulin dosage adjustment.

Safety and reliability of advice from new computer systems should be confirmed before embarking on prospective hospital trials. This process of preliminary testing is termed 'validation'. Though it forms a fundamental stage in system development, few standards exist for choosing and implementing tests. In the present paper, a validation methodology is developed in the domain of diabetes and intended for general use in chronic health management. It is based on a peer review protocol and incorporates empirical measures indicating: applicability of results to the real environment; variation among doctors; comparisons between doctors' and computer advice; and relative merits of different computer algorithms.

Algorithms

[Bacterio-expert: an integrated system for assisting in the validation of antibiotic sensitivity tests. Retrospective application in 4053 Staphylococcus].

Bacterio-expert is a simple expert system for assisting in the validation of antibiotic sensitivity testing. This system is incorporated in a data acquisition and editing program for bacteriologic test (Bacterio program written in Turbo-Pascal for personal computer users by the same authors). The principles of this system are explained and results with 4,053 antibiotic sensitivity tests on Staphylococcus aureus isolates are reported. Approximately 10% of tests required corrections.

Anti-Bacterial Agents

Validation of a detailed computer model for the electric fields in the brain.

A computer model has been designed for the calculation of the electrical fields in the head, based on the finite difference method. This method has not previously been applied for head modelling. The model was validated by using three concentric spheres and comparing it with an analytic model. Three levels of accuracy were tested. The forward solutions show that the finite difference algorithm works correctly and, by selecting the size of the volume elements properly, accurate results are obtained. The model will be applied to accurate and realistic geometries of the human head obtained from magnetic resonance images.

Brain

Measuring resource use in the ICU with computerized therapeutic intervention scoring system-based data.

BACKGROUND AND OBJECTIVE: In this era of health-care reform, there is increasing need to monitor and control health-care resource consumption. This requires the development of measurement tools that are practical, uniform, reproducible, and of sufficient detail to allow comparison among institutions, among select groups of patients, and among individual patients. We explored the feasibility of generating an index of resource use based on the Therapeutic Intervention Scoring System (TISS) from hospital electronic billing data. Such an index is potentially comparable across institutions, allows assessment of care at many levels, is well understood by clinicians, and captures many of the resources relevant to the ICU. DESIGN: We developed an automated mapping of the hospital billing database into the different items of TISS and generated computerized active TISS scores on 1,372 ICU days. The computerized score was then validated by comparison to prospectively gathered active TISS scores by trained data collectors. SETTING: Eight ICUs within a university teaching institution. PATIENTS: We studied 1,229 general medical and surgical ICU patients. INTERVENTIONS: None. MEASUREMENTS AND MAIN RESULTS: Active TISS scores ranged from 0 to 31 points. The two scores were well correlated (R2=0.53) and highly calibrated (as assessed by regression of active TISS on mean computerized active TISS [R2=0.85]). The scores were identical on 756 days (55.6%) and differed by < or = 3 TISS points on an additional 387 (28.2%) days. Interreliability assessment suggested substantial agreement (kappa statistic=0.71). The discriminatory power of the computerized score to identify different levels of ICU resource use was excellent as assessed by area under the receiver operating characteristics curves at four threshold points (0.91, 0.87, 0.89, and 0.88). Performance of the computerized score was similar across medical, coronary, and surgical ICU patient groups. CONCLUSION: An automated algorithm can reproduce valid TISS scores from standard hospital billing data, allowing comparison of patients and groups of patients in order to better understand ICU resource use.

Accounting

Development of a clinical chart to compute different disease activity indices for systemic lupus erythematosus.

Between 1990 and 1995 a European Consensus Group carried out a multicenter study to reach agreement of the definition of disease activity in systemic lupus erythematosus (SLE). A new index, the European Consensus Lupus Activity Measurement (ECLAM) index, was developed. In a second phase of the study, a prospective survey aimed at validating ECLAM and 4 other scales as steady-state and transition indices for disease activity in SLE was completed. We present the results of this survey. A standardized clinical chart was developed, together with a computer program that could automatically calculate the ECLAM score, as well as the scores for some of the disease activity scales most widely used at present, i.e., the British Isles Lupus Assessment Group, Systemic Lupus Activity Measure, SLE Disease Activity Index, and the SLE Index Score (SIS). With the participation of 28 centers in 15 different European countries, data from 121 prospectively selected new lupus patients were collected. The validity of the 5 activity scales was assessed by comparing the computed scores for each patient to a gold standard, i.e., the physician's subjective judgment on disease activity measured using a semiquantitative scale. All the indices were found to be valid instruments for measuring disease activity in SLE in both the steady-state and transition phases. The results for the various indices closely correlated with one another. Thus, the computerized chart developed by the European Consensus Group offers a simple and reliable instrument to assess disease activity and could be used to monitor lupus patients both in clinical practice and in clinical trials.

Computer Simulation

A performance evaluation of the expert system ANEMIA.

This paper reports the results of an evaluation study of the current level of performance given by ANEMIA, a knowledge-based consultation system addressing the clinical problem of managing anemic patients. ANEMIA was developed on a mainframe using the AI programming scheme EXPERT and then translated into a version running on a personal computer. At present the system is able to provide assistance in the diagnosis and management of 65 disease entities. After extensive local testing of accuracy, completeness, and consistency of the knowledge base included into ANEMIA, we designed a study to evaluate whether the system is able to appropriately mirror also the reasoning of well-known hematologists other than those who provided the knowledge. We were also interested in testing whether there were conflicting opinions among hematologists. Thus, we designed a validation study in which ANEMIA's performance could be compared with that of six hematologists and the interexpert consensus evaluated. ANEMIA's overall performance was judged acceptable in 87% (26/30) of the cases, while expert evaluators agreed with their colleagues in 90% (27/30) of them. A low interexpert consensus was found: considering the ratings given by different hematologists to the same ANEMIA performance, complete agreement occurred only 47% of the time.

Adult

Bioassay from two parabolas.

The paper deals with potency ratio estimation of parallel curve analytic dilution assays in case log dose-response relationship could be reasonably described by a parabola. It comprises: (1) testing the adequacy and validity of the quadratic model by the analysis of variance; and (2) estimation of the relative potency of an unknown in relation to a standard preparation, its standard deviation and fiducial limits for its true value. The method is applicable whenever successive doses of an unknown are a constant multiple of a standard in randomized blocks or completely randomized design. The method can be generalized to polynomial models of higher order.

Analysis of Variance

GeneGenerator--a flexible algorithm for gene prediction and its application to maize sequences.

MOTIVATION: We developed GeneGenerator because of the need for a tool to predict gene structure without knowing in advance how to score potential exons and introns in order to obtain the best results, pertinent in particular to less well-studied organisms for which suitable training sets are small. GeneGenerator is a very flexible algorithm which for a given genomic sequence generates a number of feasible gene structures satisfying user-defined constraints. The specific implementation described in detail requires minimum scoring for translation start and donor and acceptor splice sites according to previously trained logitlinear models. In addition, potential exons and introns are required to exceed specified minimal lengths and threshold scores for coding or non-coding potential derived as log-likelihood ratios of appropriate Markov sequence models. RESULTS: A database of 46 non-redundant genomic sequences from maize is used for illustration. It is shown that the correct gene structures do not always maximize the considered target function. However, in most cases, the correct or nearly correct structures are found in a small set of high-scoring structures. A critical review of the generated structures sometimes allows the choices to be narrowed by considering additional variables such as predicted splice site strength or local optimality of splice site scores. Summary statistics for prediction accuracy over all 46 maize genes are derived under cross-validation and non-cross-validation training conditions for the Markov sequence models. The algorithm achieved exon sensitivity of 0.81 and specificity of 0.75 on an independent set of 14 novel maize genomic segments. AVAILABILITY: GeneGenerator runs under Borland-Pascal 7.0 using MS-DOS and C on UNIX work stations. The source code is available upon request. CONTACT: jkleffe@euler.grumed.fu-berlin-de

Algorithms

An expert system for the analysis and interpretation of evoked potentials based on fuzzy classification: application to brainstem auditory evoked potentials.

EPEXS is an expert system for evoked potential analysis and interpretation (a medical examination performed in clinical neurophysiology laboratories), working from available clinical records and numerical data extracted from evoked potential traces. EPEXS integrates two formalisms of knowledge representation: rules and structured objects. The rules represent the elementary concepts (shallow knowledge) and include a model of possibility based on the Dubois and Prade default reasoning and possibility theory. The structured objects (prototypes) are organized as hierarchical taxonomies (underlying knowledge). These allow the description of both the objects and their relationships. The heuristics used to interpret knowledge are based on two hypotheses: the unicity of the pathological process leading to several given symptoms and the progression from the general to the specific, leading to the adoption or rejection of a class of diagnoses. This avoids the problem of the differential diagnosis. These sources of knowledge are used in a dynamical way that could be described as a four-step process: acquisition of clinical data in order to define the nosological frame of the pathology, production of hypotheses about the nature and topography of lesions, interpretation of data in accordance with these hypotheses, and finally evaluation of their likelihood. The validation shows that EPEXS topographic diagnoses were correct in 100% of cases and 92% of it nosologic diagnoses were correct, and no pathological record was interpreted as normal. When examined on a given pathology basis EPEXS was not significantly different from human experts as regards to performance, specificity, and sensitivity.

Computer Simulation

An algorithm for measurement of expiratory flow rate parameters on the partial expiratory flow-volume curve.

Partial expiratory flow-volume (PEFV) curves are a useful tool in airway challenge studies, but unlike the maximal expiratory flow-volume (MEFV) curve, lung function parameters require manual calculation from the flow-volume tracing. We describe an algorithm written in QuickBASIC that analyzes a PEFV curve superimposed on a MEFV curve by (1) identifying the PEFV curve, (2) locating the maximal expiratory flow at the point on the PEFV curve that corresponds to 60% of the baseline forced vital capacity (FVC) below total lung capacity (TLC), termed MEF40%(P), and (3) identifying the size of the PEFV curve along the TLC axis. A report of these parameters is also provided. This algorithm was validated using flow-volume curves from a clinical study in which eight subjects performed two sets of MEFV and PEFV curves separated by approximately 1 hr. Paired comparison of MEF40%(P) determined by the algorithm and two independent manual calculations correlated strongly and yielded no statistically significant differences between the two methods. We conclude that this algorithm provides rapid and accurate determinations of PEFV parameters.

Algorithms

A computer program linking physiologically based pharmacokinetic model with cancer risk assessment for breast-fed infants.

The risk assessment process predicts the chances of adverse health effects that the toxicant possibly can do to the target organism under expected conditions of exposure. Regulators chose among several mathematical approaches to estimate the risk, but in each case it is necessary to link the dosemetrics of the toxicant with its predicted health effect. In this paper, a computer program is described that allowed us to link a physiologically based pharmacokinetic (PBPK) model for tetrachloroethylene (PCE) in the lactating mother with the estimate of extra cancer risk for breast-fed infants, according to the U.S. Environmental Protection Agency (EPA) methodology. When inhaled by a lactating woman, PCE may partition into breast milk and may be transferred to the breast-fed infant. We have developed and validated experimentally a PBPK model for lactational transfer of PCE in rats, including a quantitative description of a milk compartment and the nursing pup. Subsequently, the model has been scaled to describe human physiology, and was validated with literature data for human cases of PCE exposure. Finally, we linked the dosage predictions of the PBPK model with equations used by EPA to estimate the cancer risk from PCE. The model predictions are in good agreement with both the measured values and those reported in the literature for exposure to PCE. This comparison confirms the usefulness of PBPK modeling in risk assessments.

Air Pollutants

Computerized scoring of abnormal human sleep: a validation.

A computerized assessment of sleep staging, arousals, premature ventricular contractions (PVCs), and respiratory events in sleep, was developed. Performance of the computerized system was assessed using epoch-by-epoch comparison and two human scorers across 30 consecutive patients. Percentages of agreement and Cohen's kappa coefficients were used for comparison. All agreements between all scorers for sleep staging, arousals, PVCs and respiratory events in sleep were significant (p < 0.001). The ratios of computer-human agreement descriptors to human-human agreement descriptors indicate that computerized analysis of abnormal human sleep offers reasonable results with savings in technologist time and work, but not in physician time and work.

Electroencephalography

A computer-based interview system for patients with back pain. A validation study.

A microcomputer-based system has been designed to interview patients with a view to investigating and establishing common syndromes of back and leg pain. In a randomized crossover validation study, 50 consecutive outpatients were interviewed by the computer and had a conventional clerking by a doctor. The conventional clerking made minor errors in 3.75% of questions answered and major errors in 0.90%. The computer made minor errors in 6.75% of questions and major errors in 5.45%. The majority of the computer errors were due to inadequate question design. These have been corrected, and it is anticipated that the computer will now have an overall rate of 94% correct answers and be sufficiently accurate to pursue the aim of clinical syndrome identification.

Back Pain

Computer analysis of monophasic action potentials: manual validation and clinically pertinent applications.

Monophasic action potential (MAP) recordings are increasingly being used in a variety of clinical and experimental situations but their manual measurement is cumbersome, especially when hundreds or thousands of beats must be analyzed to monitor the exact time course of action potential duration (APD) changes following heart rate alterations, during surveillance of APD alternans, or during the onset and stabilization of Class III drug effects. To facilitate this task we developed a computer program that automates programmed electrical stimulation, digitizes at 1-kHz sampling frequency MAP recordings up to 8 channels simultaneously, analyzes all APDs at repolarization levels from 10%-90% in 10% decrements (APD10-90), and automatically outputs the analyzed numerical data into spreadsheets for graphical display or statistical analysis. To validate the computer algorithm, two independent observers manually analyzed 585 concurrent MAP recordings at a paper speed of 100 mm/s. Cycle length measurements by the computer were precise to 0.4 +/- 0.5 ms as compared to the computer determined paced cycle length. Computer measurements of APD20, 50, and 90 differed from manual measurements by 2.0 +/- 8.8 ms, 0.7 +/- 7.9 ms, and 0.2 +/- 8.5 ms, respectively, for observer 1; and by 12.2 +/- 8.3 ms, 5.8 +/- 7.5 ms, and 1.4 +/- 10.1 ms, respectively, for observer 2. Inter-observer variability (IOV) was 10.3 +/- 11.1 (APD20), 5.1 +/- 9.0 ms (APD50), and 1.2 +/- 7.8 ms (APD90), which was similar to computer/observer-2 differences and significantly greater (0.001) than computer/observer-1 differences. This indicates that the computer analysis was at least as precise as manual measurements when compared to IOV, and more precise when comparing computer/observer-1 differences to IOV. While providing equal or greater precision, computer-aided analysis of 100 MAP signals took approximately 1 minute while manual analysis of the same data set took between 2.5 and 4 hours. The pacing and analysis software was subsequently applied to experiments that mimic clinically pertinent examples of MAP recordings: (1) automatic generation, analysis, and graphical display of electrical restitution curves at multiple ventricular sites simultaneously; (2) evaluation of myocardial pharmacokinetics by monitoring the progression of Class III antiarrhythmic drug effects by continuous MAP recordings, and displaying differences in drug action between multiple sites; (3) depiction of the adaptation time course of APD to abrupt changes in paced cycle length; and (4) quantitative analysis of APD alternans during myocardial ischemia. The results show that our computerized algorithm greatly facilitates the generation of cardiac electrophysiological, and clinically important, data.

Action Potentials

Validating a decision support system for anti-epileptic drug treatment. Part II: adjusting anti-epileptic drug treatment.

A model of expertise for monitoring antiepileptic drug treatment was implemented in a decision support system. We validated the advice of the system regarding treatment decisions at first follow-up with 265 paper cases based on patient records. The reference for comparison is based on the opinions of neurologists. We found considerable variation among the decisions of five neurologists. It could be shown that the system agreed with (groups of) neurologists at least as often as individual neurologists did. The correctness of the system was consistently higher than that of each of the neurologists, when the majority decision of the remaining neurologists constituted the standard.

Anticonvulsants

Detecting errors in a scoring program: a method of double diagnosis using a computer-generated sample.

This paper discusses a new method for locating errors in diagnostic computer scoring programs for structured clinical interviews. It was proposed as a test of the accuracy of the scoring program for the Composite International Diagnostic Interview, version 1.1. The proposal was to create an independent scoring program in a different computer language but serving the same criteria. Both programs were then applied to the same large set of valid (i.e., logically consistent) computer-generated test cases, and differences in diagnostic assignments reviewed. The method described can identify the program steps that account for the sources of the errors. Corrections can be made and the programs run again on new sets of test cases until discrepancy-free results are achieved. While this method cannot discover errors that are repeated in the two programs, it does discover more of the errors in a scoring program than we have previously been able to identify. This technique provides a systematic and rigorous approach to assuring the accuracy of scoring programs based on established algorithms.

Algorithms

Machine learning for an expert system to predict preterm birth risk.

OBJECTIVE: Develop a prototype expert system for preterm birth risk assessment of pregnant women. Normal gestation involves a term of 40 weeks, but because 8-12% of the newborns in the United States are delivered prior to 37 weeks' gestation, problems associated with prematurity continue to plague individuals, families, and the health care system. DESIGN: A knowledge-base development methodology used machine learning, statistical analysis, and validation techniques to analyze three large datasets (18,890 subjects and 214 variables). The dependent (i.e., decision) variable studied was weeks of gestation at delivery, with dichotomous coding of preterm delivery (prior to 37 weeks) and full-term delivery (37+ weeks). RESULTS: Machine learning with a program named Learning from Examples using Rough Sets (LERS) induced 520 usable rules that were entered into a prototype expert system. The prototype expert system was 53-88% accurate in predicting preterm delivery for 9,419 patients. CONCLUSION: The prototype expert system was more accurate than traditional manual techniques in predicting preterm birth.

Adult

Low-energy imaging with high-energy bremsstrahlung beams: analysis and scatter reduction.

The contrast and zero spatial frequency signal-to-noise ratio produced by a method for radiation therapy portal imaging known as low-energy imaging with high-energy bremsstrahlung beams have been mathematically analyzed. The analysis makes extensive use of Monte Carlo techniques and incorporates the detector, the spectrum, phantom, and geometry. The analysis is validated through comparison with measured data including subject contrast measurements and the attenuation of the beam with lead. Scatter reduction is found to be potentially the most effective method to improve contrast and SNR for a film based system. A large fraction of the scatter detected is of a much higher energy than that found in diagnostic radiology. Hence, traditional antiscatter grids, such as those used in diagnostic radiology, are ineffective. The analysis and theory from the literature are applied to design a new grid which is more appropriate for this application. The grid produces a modest improvement according to a contrast-detail study.

Humans