PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Software Validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Validation, clinical trial, and evaluation of a radiology expert system.

The PHOENIX Radiology Consultant is a rule-based expert system which assists physicians in planning radiological work-up strategies. This article describes the methods used to create and validate the system's knowledge base. The feasibility and acceptability of PHOENIX were tested for two years in a clinical trial. During this period, the system was used 1,421 times, an average of 13.7 times per week, primarily by medical students and nonradiologist physicians. Much of the system's use occurred at night and on weekends, when the radiology department was not fully staffed. Several physicians were enlisted to further evaluate the utility of the system. The results of their evaluation indicate that an expert system that helps physicians select diagnostic-imaging studies can serve as a useful and informative component of a radiology information system, and is particularly useful for medical students and physicians in training.

Algorithms↗

TopFit: a PC-based pharmacokinetic/pharmacodynamic data analysis program.

The program TopFit was developed and validated within the European pharmaceutical industry. It provides both pharmacokinetic data analysis support for international regulatory submissions of new drugs, and sophisticated techniques for model-based kinetic/dynamic evaluation during drug development. TopFit features are: (1) non-compartmental methods; (2) standard compartment models assembled from input and disposition modules; (3) a potentially unlimited number of linear user-defined models that accommodate metabolites, effects, and absorption profiles; (4) a library of 24 non-linear models. No user programming is required. A well-defined file structure allows ready exchange of data with other programs such as SAS. TopFit version 2.0 is now commercially available, with comprehensive documentation, in the form of an MS-DOS application for the PC.

Data Interpretation, Statistical↗

Geriatric patient simulations for dental hygiene.

The rapidly increasing number of this country's elderly requires that dental hygiene students practice the clinical problem-solving skills of information gathering, assessment, and treatment applied to geriatric patients. Computer-based simulations are purported to provide this experience, but little research has been completed with simulations in the education of dental hygienists. This paper summarizes the process used to design, develop, and evaluate a series of eighteen computer-based geriatric simulated patients. It contains a brief description of the simulations and a description of the design, validation, authoring, and formative evaluation phases. The paper also describes the summative evaluation, provides implementation suggestions, and summarizes future directions. The summative evaluation, conducted at four institutions, suggests that computer-based simulations are an effective instructional method as measured by pre/post-tests. The results suggest that simulations can provide a standardized set of geriatric patient experiences. These simulations may prove especially valuable at institutions that are unable to provide clinical geriatric experiences or lack the expertise to conduct a didactic course in geriatrics.

Aged↗

Medical data and knowledge management by integrated medical workstations: summary and recommendations.

The health care professional workstation will function as an interface between the user and the patient data as well as an interface pertinent medical knowledge. Appropriate knowledge focus will require the workstation to recognize the concepts and structure of patient data, and understand the scope and access methods of knowledge sources. Issues are organized around five major themes: (i) structure, (ii) reliability and validation, (iii) views, (iv) location, and (v) ethical and legal. Conventional database representations can effectively address data structure and format variations that will inevitably persist in local data stores. The reliability of data and the validation of knowledge are critical issues that may determine the ultimate utility of clinical workstations. Alternative views of patient information and knowledge sources represent the true power of an intelligent data portal, represented by a well-designed clinical workstation. Both data and knowledge are optimally represented in decentralized information networks, although the confidentiality and ownership of this information must be respected. Evolutionary progress toward consistent representations of knowledge and patient data will be facilitated by the establishment of self-documentation standards for the developers of data encoding systems and knowledge sources, perhaps extended from the preliminary model afforded by the Unified Medical Language System (UMLS).

Computer Security↗

Reproducibility of polar map generation and assessment of defect severity and extent assessment in myocardial perfusion imaging using positron emission tomography.

The purpose of this study was to determine the reliability of new software developed for the analysis of cardiac tomographic data. The algorithm delineates the long axis and defines the basal plane and subsequently generates polar maps to quantitatively and reproducibly assess the size and severity of perfusion defects. The developed technique requires an initial manual estimate of the left ventricular long axis and calculates the volumetric maximum myocardial activity distribution. This surface is used to map three-dimensional tracer accumulation onto a two-dimensional representation (polar map), which is the basis for further processing. The spatial information is used to compute geometrical and mechanical properties of a solid model of the left ventricle including the left heart chamber. A new estimate of the axis is determined from this model, and the previously outlined procedure is repeated together with an automated definition of the valve plane until differences between the polar maps can be neglected. This quantitative analysis software was validated in phantom studies with defects of known masses and in ten data sets from normals and patients with coronary artery disease of various severity. We investigated the reproducibility of the maps with the introduction of a similarity criterion where the ratio of two corresponding polar map elements lies within a 10% interval. The maps were also used to measure intra-and interobserver variability in respect of defect size and severity. In the phantom studies, it was possible to reliably assess mass information over a wide range of defects from 5 to 60 g (slope: 1.02, offset -0.68, r = 0.972). Patient studies revealed a statistically significant increase in the reproducibility of the automatic technique compared with the manual approach: 54%+/-19% (manual) compared with 88%+/-9% (automatic) for observer 1 and 61%+/-20% vs 82%+/-5% for observer 2, respectively. The intervariability analysis showed a significant improvement from 59%+/-14% to 83%+/-7% in similar polar map elements and a significantly improved correlation in the calculation of severity (from r = 0.908 to 0.989) and extent (from r = 0.963 to r = 0.992) of the perfusion defects when the automated procedure was applied. It is concluded that, assuming a constant wall thickness and tissue density, absolute defect mass can be reliably estimated. Furthermore, the proposed software demonstrates a significant improvement in the generation of volumetric polar maps for the quantitative assessment of perfusion defects.

Algorithms↗

An automatic approach to the analysis, quantitation and review of perfusion and function from myocardial perfusion SPECT images.

UNLABELLED: We have developed a software suite that automatically selects, analyses, quantitates and displays all the key image data in a myocardial perfusion SPECT study. METHODS: The files automatically selected (upon specification of the patient name) are rest and stress projections, rest and stress short axis and gated short axis files, and all 'snapshot' files. The projection data sets are presented in cine mode for evaluation of patient motion, while the lung/heart ratio at rest and stress is calculated from regions of interest (ROIs) that are automatically derived and overlayed on the LAO 45 images. Left ventricular (LV) cavity volumes at rest and stress are calculated from the short axis data sets, and the related transient ischemic dilation (TID) ratio derived and displayed. Quantitative measurements of global (ejection fraction) and regional function parameters are performed from the gated short axis dataset. All algorithms use the C++, X-Windows and OSF-Motif standards. The overall suite executes in less than 1 minute on a SunSPARC5 with 32 Mb of RAM and no proprietary hardware. RESULTS: The software was validated on 144 patients (118 rest 201T1/post-stress 99mTc-sestamibi, 18 post-stress 99mTC-sestamibi, 8 rest 201Tl) acquired on a 90 degrees dual detector (ADAC Vertex, 91 patients) and a triple detector camera (Picker Prism 3000, 53 patients). Overall, the individual algorithms for the analysis of projection, short axis and gated short axis images were successful in 622/660 (94.2%) of the images. In 80.5% of the patients (73/91 + 43/53) all algorithms executed successfully, without significant difference in success rates for 201Tl versus 99mTc-sestamibi images. CONCLUSION: Our automated approach to myocardial perfusion SPECT analysis and review is highly successful, intrinsically reproducible, and can produce time and cost savings while improving accuracy in a clinical or teleradiology-type environment.

Algorithms↗

Validation of the medical expert system RENOIR.

RENOIR is an expert system developed to assist the diagnosis of 37 diseases of connective tissue and inflammatory arthropathies. Precise diagnosis of rheumatic diseases implies great uncertainty and there is no gold standard with which to compare the expert system output. To overcome this problem a set of clinical cases was submitted to RENOIR and its diagnoses were compared with those of clinicians. Medical records of 81 patients with rheumatic diseases were interpreted by RENOIR and by 12 clinicians at three different expertise levels in rheumatology. Distances between the likelihoods of the 37 considered diseases provided by clinicians and RENOIR were computed as a disagreement measure. Mahalanobis distance was used to correct the collinearity between the possibilities of each pair of diseases. Using the resulting matrices of distances between experts, cluster analyses were carried out to classify RENOIR among human experts. Greater differences between RENOIR and clinicians than among clinicians themselves were not found.

Cluster Analysis↗

Automatic detection of wave boundaries in multilead ECG signals: validation with the CSE database.

This paper presents an algorithm for automatically locating the waveform boundaries (the onsets and ends of P, QRS, and T waves) in multilead ECG signals (the 12 standard leads and the orthogonal XYZ leads). Given these locations, features of clinical importance (such as the RR interval, the PQ interval, the QRS duration, the ST segment, and the QT interval) may be measured readily. First, a multilead QRS detector locates each beat, using a differentiated and low-pass filtered ECG signal as input. Next, the waveform boundaries are located in each lead. The leads in which the detected electrical activity is of longest duration are used for the final determination of the waveform boundaries. The performance of our algorithm has been evaluated using the CSE multilead measurement database. In comparison with other algorithms tested by the CSE, our algorithm achieves better agreement with manual measurements of the T-wave end and of interval values, while its measurements of other waveform boundaries are within the range of the algorithm and manual measurements obtained by the CSE.

Algorithms↗

Analysis of brain and cerebrospinal fluid volumes with MR imaging. Part I. Methods, reliability, and validation.

A computerized system was developed to process standard spin-echo magnetic resonance (MR) imaging data for estimation of brain parenchyma and cerebrospinal fluid (CSF) volumes. In phantom experiments, the estimated volumes corresponded closely to the true volumes (r = .998), with a mean error less than 1.0 cm3 (for phantom volumes ranging from 5 to 35 cm3), with excellent intra- and interobserver reliability. In a clinical validation study with actual brain images of 10 human subjects, the average coefficient of variation between observers for the measurement of absolute brain and CSF volumes was 1.2% and 6.4%, respectively. The intraclass correlations for three expert operators is greater than .99 in the measurement of brain and ventricular volumes and greater than .94 for total CSF volume. Therefore, the authors believe that their technique to analyze MR images of the brain performed with acceptable levels of accuracy and reliability and that it can be used to measure brain and CSF volumes for clinical research. This technique could be helpful in the correlation of neuroanatomic measurements to behavioral and physiologic parameters in neuropsychiatric disorders.

Algorithms↗

Toward an intelligent wound assessment system.

There is general agreement regarding the need for pressure ulcer assessment methodology which more discretely reflects relevant aspects of wound status than does the commonly used staging system. The Pressure Sore Status Tool (PSST) is one such instrument which was developed with consensual expert input. While the psychometric properties of the PSST have been reported in the literature, the instrument was validated using ET nurses, highly trained wound care specialists, and existed only in manual form. This paper reports results from attempts to establish reliability estimates for healthcare practitioners without extraordinary wound care training or experience. The paper further describes the automation of the PSST and provides examples of pressure ulcer profiles tracked over time. Results indicate that inter-rater reliability with general healthcare practitioners was .78 and intra-rater reliability was .89. The practitioners were able to use the PSST for over six months and the automated system allowed analysis of wound healing profiles that would have been difficult using a manual system. These results imply that movement toward an automated system which makes discriminations regarding the effects of various treatment and intervention strategies is possible and practical.

Aged↗

The role of quality assurance in computer inspections.

Changing technology affords the Quality Assurance auditor with the challenge of applying computer validation concepts to a variety of computer system types. In addition, these technology changes have caused the developers role to change as well. In an innovative research facility, the developer may include an in-house professional group, a vendor, or an end-user. With these issues in mind, the QA auditor needs a tool to accomplish the task of inspecting systems as they are being created. The prospective inspection process is the tool for accomplishing this task. This inspection involves the QA auditor's involvement in the development of a new system, as a member of the development team, from the initial creation through the implementation of the computer system. This presentation will focus on illustrating the steps in conducting the prospective inspection process, from expected deliverables and document reviews to final report and management notification. The benefits of QA involvement in the development process will also be discussed.

Clinical Laboratory Information Systems↗

The performance of the knowledge-based system VALAB revisited: an evaluation after five years.

In 1988, inundated by the tedious work of validation of laboratory reports in a large hospital biochemistry laboratory, we designed VALAB, a knowledge-based system specially dedicated to this iterative function. Coping at first with a few biochemical tests, the program has been progressively expanded to forty-five common chemical tests. Simultaneously some new rules have been introduced to "weight" the conclusion in different circumstances and rules taking into consideration some clinical data have also been written. Moreover the program moved to other disciplines, pH and blood gases, haematology and coagulation. Accordingly the evaluation protocol has been modified, incorporating a new step, the consensus decision of the pathologists, operating within the initial protocol and based upon the various criteria of epidemiology. These major changes and improvements have led us to check and describe again the performance of this updated VALAB knowledge-based system.

Artificial Intelligence↗

Computerized decision support for concurrent utilization review using the HELP system.

OBJECTIVE: Development and evaluation of computerized concurrent utilization review (UR) support taking advantage of a clinically rich computerized patient database. DESIGN: The Automated Support System for Utilization Review (ASSURE) applies the Appropriateness Evaluation Protocol (AEP) Day of Care criteria to computerized patient data in the HELP hospital information system. This paper reports the development, verification, and validation of ASSURE. MEASUREMENTS: Implementation correctness was verified by measuring agreement with a nurse reviewer, using separate sample sets for all 20 criteria for a total of 560 current inpatients. Usefulness in detecting inappropriate days of care was validated by two nurse reviewers who were crossed with manual and computer-assisted review methods in a blocked design for 168 current inpatients. Agreement with reviewers, sensitivity, specificity, positive predictive value, and negative predictive value were measured. RESULTS: Agreement was very good for satisfaction of criteria, and good for appropriateness of day of care. A patient day identified by ASSURE as potentially inappropriate would be twice as likely to be judged inappropriate by a reviewer as a randomly selected patient day. Review of the 10% of patient days identified as potentially inappropriate by ASSURE would identify approximately 21% of the inappropriate days of care. CONCLUSION: ASSURE is a clinically useful tool for screening adult acute care patients for inappropriate days of care, and promises to make a major contribution to reducing health care costs. The prognosis for successful routine clinical use is good.

Artificial Intelligence↗

A program for the user-independent computation of the correlation dimension and the largest Lyapunov exponent of heart rate dynamics from small data sets.

We propose a specially optimized computer program for the user-independent calculation of the correlation dimension D and the largest Lyapunov exponent L of heart rate dynamics on the basis of only 1024 electrocardiographically recorded RR intervals (heartbeat intervals). The validity of our program was established by analyzing a set of artificial standard signals. Our norm values of the correlation dimension (D = 5.37 +/- 0.62) and the largest Lyapunov exponent (L = 0.561 +/- 0.037 bits/beat) of RR dynamics, obtained from 79 healthy adults aged 26.3 +/- 4.8 years, were independent of gender and age; D and L correlated slightly with each other (Pearson correlation coefficient r = 0.26). Short-term reliability, tested for 25 of our subjects by two successive recordings, was fair: the intraclass correlation coefficients (ICCs) were 0.45 and 0.41 for D and L of RR dynamics, respectively. However, long-term reliability, tested for eight of our subjects by ten weekly recordings, was acceptable for L (ICC = 0.40) but not for D (ICC = 0.01). These results permit group comparisons on the basis of single measurements of L of RR dynamics. A reliable differentiation between young healthy individuals requires four measurements of L.

Adult↗

Use of an artificial neural network (ANN) for classifying nursing care needed, using incomplete input data.

BACKGROUND: In German nursing insurance, the act of classifying the client into four categories of disability is based on legally defined distinct criteria. When classifying deceased persons it is often impossible to collect all the required information. PRIMARY OBJECTIVE: We aimed to determine the ability of an artificial neural network (ANN) to calculate the category of disability, to investigate the response of the ANN to input items of different nature, quantity and data quality, and to estimate the minimum number of training data required. RESEARCH DESIGN: The investigation was conducted as a retrospective observational study. METHODS AND PROCEDURES: The analysis was based on routine records of 14000 adult clients of the nursing insurance. Several ANNs were trained, varying nature, number and quality of the input items as well as the size of the training data set. Each ANN's classification competence was tested on independent validation data, judging the ANN's conformance to the result of the individual expert assessment, using kappa statistics. MAIN RESULTS: Fed with all 30 input items available, the net classified 80% of cases correctly (weighted kappa = 0.78). Using three input items, weighted kappa was 0.63. Severe misclassification (deviation by more than one category in either direction) ranged between 0.2% (all 30 input items) and 3.7% (3/30 items). The less complete the individual input items were, the less accurate was the net's estimate. A 20% rate of missing values was well tolerated. A training set comprising 500 cases was adequate. CONCLUSIONS: The input item set inherits redundancy. The ANN's ability to correctly respond to subsets of input items makes it a powerful tool in quality control. In the categorization of deceased persons when only an incomplete input item set is available, the ANN can achieve satisfactory results.

Activities of Daily Living↗

Monitoring expert system performance using continuous user feedback.

OBJECTIVE: To evaluate the applicability of metrics collected during routine use to monitor the performance of a deployed expert system. METHODS: Two extensive formal evaluations of the GermWatcher (Washington University School of Medicine) expert system were performed approximately six months apart. Deficiencies noted during the first evaluation were corrected via a series of interim changes to the expert system rules, even though the expert system was in routine use. As part of their daily work routine, infection control nurses reviewed expert system output and changed the output results with which they disagreed. The rate of nurse disagreement with expert system output was used as an indirect or surrogate metric of expert system performance between formal evaluations. The results of the second evaluation were used to validate the disagreement rate as an indirect performance measure. Based on continued monitoring of user feedback, expert system changes incorporated after the second formal evaluation have resulted in additional improvements in performance. RESULTS: The rate of nurse disagreement with GermWatcher output decreased consistently after each change to the program. The second formal evaluation confirmed a marked improvement in the program's performance, justifying the use of the nurses' disagreement rate as an indirect performance metric. CONCLUSIONS: Metrics collected during the routine use of the GermWatcher expert system can be used to monitor the performance of the expert system. The impact of improvements to the program can be followed using continuous user feedback without requiring extensive formal evaluations after each modification. When possible, the design of an expert system should incorporate measures of system performance that can be collected and monitored during the routine use of the system.

Expert Systems↗

SuperStar: a knowledge-based approach for identifying interaction sites in proteins.

An empirical method for identifying interaction sites in proteins is described and validated. The method is based entirely on experimental information about non-bonded interactions occurring in small-molecule crystal structures. These data are used in the form of scatterplots that show the experimentally observed distribution of one functional group (the "contact group" or "probe") around another. A template molecule (e.g. a protein binding site) is broken down into structure fragments and the scatterplots, showing the distribution of a chosen probe around these structure fragments, are superimposed on the corresponding parts of the template. The scatterplots are then translated into a three-dimensional map that shows the propensity of the probe at different positions around the template molecule. The method is illustrated for l -arabinose-binding protein, complexed with l -arabinose and with d -fucose, and for dihydrofolate reductase complexed with methotrexate. The method is validated on 122 X-ray structures of protein-ligand complexes. For all the binding sites of these proteins, propensity maps are generated for four different probes: a charged NH+3nitrogen, a carbonyl oxygen, a hydroxyl oxygen and a methyl carbon atom. Next, the maps are compared with the experimentally observed positions of ligand atoms of these types. For 74% of these ligand atoms (84% of the solvent-inaccessible ones) the calculated propensity of the matching probe at the experimental positions is higher than expected by chance. For 68% of the atoms (82% of the solvent-inaccessible ones) the propensity of the matching probe is higher than that of the other three probes. These results indicate that the approach generally gives good predictions for protein-ligand interactions. The potential applications of the propensity maps range from an aid in manual docking and structure-based drug design to their use in pharmacophore development.

Artificial Intelligence↗

Mean regional cerebral blood flow images of normal subjects using technetium-99m-HMPAO by automated image registration.

UNLABELLED: The purpose of this study was twofold: to calculate relative uptake values for 99mTc-HMPAO in various regions of the normal brain after alignment and registration to a standard shape and size, and to validate the automated image registration (AIR) program for SPECT-to-SPECT transformation. METHODS: Thirty subjects took part in this study. Technetium-99m-HMPAO brain SPECT and x-ray-CT scans were acquired. SPECT images were normalized to an average activity of 100 counts/pixel. Intersubject accuracy was evaluated on brain images of 17 normal subjects (mean age = 64.9 +/- 8.7 yr). These images were aligned and registered to a standard size and shape with the help of AIR. Realigned images were overlaid on reference images to determine the overlap areas. Intrasubject accuracy was evaluated by realigning 20 degree rotated brain images with an index calculated as: overlap area/(overlap area + nonoverlap area). Anatomical variability between realigned target and reference images was evaluated by measurements on corresponding x-ray-CT scans, realigned using transformations that were established by the SPECT images. Realigned brain SPECT images of 30 normal subjects (mean age = 50.7 +/- 18.7 yr), including those subjects examined in the accuracy validation study, were used to generate mean and s.d. images. Images based on the mean value of each voxel (n = 30) were compared with other mean images prepared by the human brain atlas (HBA) standardization technique on a voxel-by-voxel basis to generate T maps. RESULTS: Accuracy indices were 0.98 +/- 0.006 and 0.99 +/- 0.002 for the intersubject and intrasubject evaluations, respectively. The maximum anatomical variability was 4.7 mm after realignment. Paired Student's t-test comparisons of mean HBA and AIR images revealed statistically significant differences for the deep white matter, pons and occipito-temporal regions. These differences could be explained by variation in the population being studied and the protocol for data handling by AIR and HBA. CONCLUSION: AIR aligns and registers brain SPECT images with acceptable accuracy, without the necessity of MRI or x-ray-CT scans.

Brain↗