PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Boosting regression estimators.

There is interest in extending the boosting algorithm (Schapire, 1990) to fit a wide range of regression problems. The threshold-based boosting algorithm for regression used an analogy between classification errors and big errors in regression. We focus on the practical aspects of this algorithm and compare it to other attempts to extend boosting to regression. The practical capabilities of this model are demonstrated on the laser data from the Santa Fe times-series competition and the Mackey-Glass time series, where the results surpass those of standard ensemble average.

Algorithms↗

Management of epilepsy in adults. Diagnosis guidelines.

In this first of two articles on new epilepsy guidelines for primary care physicians, the authors present detailed algorithms for the diagnosis and classification of seizure disorders in adults. They discuss the differentiation between generalized and partial seizures and stress that accurate identification is especially important because the type of seizure determines the appropriate treatment. The second article (page 29) looks at the treatment portion of the new guidelines.

Adolescent↗

Distinction of brain tissue, low grade and high grade glioma with time-resolved fluorescence spectroscopy.

Neuropathology frozen section diagnoses are difficult in part because of the small tissue samples and the paucity of adjunctive rapid intraoperative stains. This study aims to explore the use of time-resolved laser-induced fluorescence spectroscopy as a rapid adjunctive tool for the diagnosis of glioma specimens and for distinction of glioma from normal tissues intraoperatively. Ten low grade gliomas, 15 high grade gliomas without necrosis, 6 high grade gliomas with necrosis and/or radiation effect, and 14 histologically uninvolved "normal" brain specimens are spectroscopicaly analyzed and contrasted. Tissue autofluorescence was induced with a pulsed Nitrogen laser (337 nm, 1.2 ns) and the transient intensity decay profiles were recorded in the 370-500 nm spectral range with a fast digitized (0.2 ns time resolution). Spectral intensities and time-dependent parameters derived from the time-resolved spectra of each site were used for tissue characterization. A linear discriminant analysis diagnostic algorithm was used for tissue classification. Both low and high grade gliomas can be distinguished from histologically uninvolved cerebral cortex and white matter with high accuracy (above 90%). In addition, the presence or absence of treatment effect and/or necrosis can be identified in high grade gliomas. Taking advantage of tissue autofluorescence, this technique facilitates a direct and rapid investigation of surgically obtained tissue.

Algorithms↗

ROC and CART analysis of subcutaneous adipose tissue topography (SAT-Top) in type-2 diabetic women and healthy females.

Women suffering from type-2 diabetes mellitus (NIDDM) show a more android fat pattern than healthy females, but to date no exact determination of their fat distribution differences exists. Measurements at 15 specified body sites with an optical device, the LIPOMETER, provide a subcutaneous adipose tissue topography (SAT-Top) of the individual. SAT-Top of 20 female NIDDM patients and 122 healthy controls was measured. ROC curve analysis was applied to evaluate the discriminative power of each body site and to provide cutoff values. Then a classification tree by the CART algorithm was established, showing SAT-Top differences between the two groups. Best discriminating results were achieved by the neck site (ROC area index = 0.76, sensitivity = 61.3%, specificity = 77.8%), the four sites of the thigh (area indices from 0.71 to 0.76), and a linear combination of all body sites stemming from a previous factor analysis, which provides condensed information of the extremities SAT-Top (area index = 0.80, sensitivity = 80.4%, specificity = 64.6%). The results could be improved by a summary measure of "android fat pattern" (area index = 0.89, sensitivity = 73.6%, specificity =88.3%) and a proportional measure of SAT-distribution, the relative neck (area index = 0.84, sensitivity = 83.0%, specificity = 70.5%). Overall, 136 (95.8%) of the 142 subjects were correctly classified by the classification tree (sensitivity = 75%, specificity = 99.2%). Both methods show the expected increased upper trunk obesity and decreased lower body obesity of NIDDM women compared with healthy females. Am. J. Hum. Biol. 12:388-394, 2000. Copyright 2000 Wiley-Liss, Inc.

Journal Article↗

Diagnosis of meningioma by time-resolved fluorescence spectroscopy.

We investigate the use of time-resolved laser-induced fluorescence spectroscopy (TR-LIFS) as an adjunctive tool for the intraoperative rapid evaluation of tumor specimens and delineation of tumor from surrounding normal tissue. Tissue autofluorescence is induced with a pulsed nitrogen laser (337 nm, 1.2 ns) and the intensity decay profiles are recorded in the 370 to 500 nm spectral range with a fast digitizer (0.2 ns resolution). Experiments are conducted on excised specimens (meningioma, dura mater, cerebral cortex) from 26 patients (97 sites). Spectral intensities and time-dependent parameters derived from the time-resolved spectra of each site are used for tissue characterization. A linear discriminant analysis algorithm is used for tissue classification. Our results reveal that meningioma is characterized by unique fluorescence characteristics that enable discrimination of tumor from normal tissue with high sensitivity (>89%) and specificity (100%). The accuracy of classification is found to increase (92.8% cases in the training set and 91.8% in the cross-validated set correctly classified) when parameters from both the spectral and the time domain are used for discrimination. Our findings establish the feasibility of using TR-LIFS as a tool for the identification of meningiomas and enables further development of real-time diagnostic tools for analyzing surgical tissue specimens of meningioma or other brain tumors.

Algorithms↗

A tutorial on the use of ROC analysis for computer-aided diagnostic systems.

The application of the receiver operating characteristic (ROC) curve for computer-aided diagnostic systems is reviewed. A statistical framework is presented and different methods of evaluating the classification performance of computer-aided diagnostic systems, and, in particular, systems for ultrasonic tissue characterization, are derived. Most classifiers that are used today are dependent on a separation threshold, which can be chosen freely in many cases. The separation threshold separates the range of output values of the classification system into different target groups, thus conducting the actual classification process. In the first part of this paper, threshold specific performance measures, e.g., sensitivity and specificity, are presented. In the second part, a threshold-independent performance measure, the area under the ROC curve, is reviewed. Only the use of separation threshold-independent performance measures provides classification results that are overall representative for computer-aided diagnostic systems. The following text was motivated by the lack of a complete and definite discussion of the underlying subject in available textbooks, references and publications. Most manuscripts published so far address the theme of performance evaluation using ROC analysis in a manner too general to be practical for everyday use in the development of computer-aided diagnostic systems. Nowadays, the user of computer-aided diagnostic systems typically handles huge amounts of numerical data, not always distributed normally. Many assumptions made in more or less theoretical works on ROC analysis are no longer valid for real-life data. The paper aims at closing the gap between theoretical works and real-life data. The review provides the interested scientist with information needed to conduct ROC analysis and to integrate algorithms performing ROC analysis into classification systems while understanding the basic principles of classification.

Diagnosis, Computer-Assisted↗

Feature selection for optimized skin tumor recognition using genetic algorithms.

In this paper, a new approach to computer supported diagnosis of skin tumors in dermatology is presented. High resolution skin surface profiles are analyzed to recognize malignant melanomas and nevocytic nevi (moles), automatically. In the first step, several types of features are extracted by 2D image analysis methods characterizing the structure of skin surface profiles: texture features based on cooccurrence matrices, Fourier features and fractal features. Then, feature selection algorithms are applied to determine suitable feature subsets for the recognition process. Feature selection is described as an optimization problem and several approaches including heuristic strategies, greedy and genetic algorithms are compared. As quality measure for feature subsets, the classification rate of the nearest neighbor classifier computed with the leaving-one-out method is used. Genetic algorithms show the best results. Finally, neural networks with error back-propagation as learning paradigm are trained using the selected feature sets. Different network topologies, learning parameters and pruning algorithms are investigated to optimize the classification performance of the neural classifiers. With the optimized recognition system a classification performance of 97.7% is achieved.

Algorithms↗

Learning weighted metrics to minimize nearest-neighbor classification error.

In order to optimize the accuracy of the Nearest-Neighbor classification rule, a weighted distance is proposed, along with algorithms to automatically learn the corresponding weights. These weights may be specific for each class and feature, for each individual prototype, or for both. The learning algorithms are derived by (approximately) minimizing the Leaving-One-Out classification error of the given training set. The proposed approach is assessed through a series of experiments with UCI/STATLOG corpora, as well as with a more specific task of text classification which entails very sparse data representation and huge dimensionality. In all these experiments, the proposed approach shows a uniformly good behavior, with results comparable to or better than state-of-the-art results published with the same data so far.

Algorithms↗

Pathologic diagnosis of acute lymphocytic leukemia.

With present knowledge, the optimal management of individual patients with acute leukemia requires that every case be studied by morphology, cytochemistry, cytogenetic, immunologic and molecular techniques. An algorithm for diagnostic evaluation and classification of ALL is provided in Fig. 11. Other techniques, such as DNA or cDNA [figure: see text] microarray, are at present important research tools but have not yet had a major effect on patient care. More detailed studies of individual patients need to be conducted at specialized cancer centers, where preservation of cells, DNA, RNA, or protein is possible. Such investigations will yield important information on the clinical importance of the expression of various markers, the prevalence and relevance of bilineage and biphenotypic leukemias, and above all will reveal the mechanisms of leukemogenesis and of disease evolution. Such insights will further aid clinicians in treating ALL and in preventing refractory disease.

Adult↗

Rapid identification of mycolic acid patterns of mycobacteria by high-performance liquid chromatography using pattern recognition software and a Mycobacterium library.

Current methods for identifying mycobacteria by high-performance liquid chromatography (HPLC) require a visual assessment of the generated chromatographic data, which often involves time-consuming hand calculations and the use of flow charts. Our laboratory has developed a personal computer-based file containing patterns of mycolic acids detected in 45 species of Mycobacterium, including both slowly and rapidly growing species, as well as Tsukamurella paurometabolum and members of the genera Corynebacterium, Nocardia, Rhodococcus, and Gordona. The library was designed to be used in conjunction with a commercially available pattern recognition software package, Pirouette (Infometrix, Seattle, Wash.). Pirouette uses the K-nearest neighbor algorithm, a similarity-based classification method, to categorize unknown samples on the basis of their multivariate proximities to samples of a preassigned category. Multivariate proximity is calculated from peak height data, while peak heights are named by retention time matching. The system was tested for accuracy by using 24 species of Mycobacterium. Of the 1,333 strains evaluated, > or = 97% were correctly identified. Identification of M. tuberculosis (n = 649) was 99.85% accurate, and identification of the M. avium complex (n = 211) was > or = 98% accurate; > or = 95% of strains of both double-cluster and single-cluster M. gordonae (n = 47) were correctly identified. This system provides a rapid, highly reliable assessment of HPLC-generated chromatographic data for the identification of mycobacteria.

Algorithms↗

Unsupervised image classification of medical ultrasound data by multiresolution elastic registration.

Thousands of medical images are saved in databases every day and the need for algorithms able to handle such data in an unsupervised manner is steadily increasing. The classification of ultrasound images is an outstandingly difficult task, due to the high noise level of these images. We present a detailed description of an algorithm based on multiscale elastic registration capable of unsupervised, landmark-free classification of cardiac ultrasound images into their respective views (apical four chamber, two chamber, parasternal long axis and short axis views). We validated the algorithm with 90 unselected, consecutive echocardiographic images recorded during daily clinical work. When the two visually very similar apical views (four chamber and two chamber) are combined into one class, we obtained a 93.0% correct classification (chi2 = 123.8, p < 0.0001, cross-validation 93.0%; chi2 = 131.1, p < 0.0001). Classification into the 4 classes reached a 90.0% correct classification (chi2 = 205.4, p < 0.0001, cross-validation 82.2%; chi2 = 165.9, p < 0.0001).

Algorithms↗

Application of K-nearest neighbors algorithm on breast cancer diagnosis problem.

This paper addresses the Breast Cancer diagnosis problem as a pattern classification problem. Specifically, this problem is studied using the Wisconsin-Madison Breast Cancer data set. The K-nearest neighbors algorithm is employed as the classifier. Conceptually and implementation-wise, the K-nearest neighbors algorithm is simpler than the other techniques that have been applied to this problem. In addition, the Knearest neighbors algorithm produces the overall classification result 1.17% better than the best result known for this problem.

Algorithms↗

Transferability of neural network-based decision support algorithms for early assessment of chest-pain patients.

The present investigation concerns methodological and epidemiological aspects of the transferability of artificial neural network-based algorithms, as key-components for classification in decision support systems (DSS). The prevalence of pathological conditions to be detected must be known in order to tune an artificial neural networks (ANN)-decision algorithm so that the predictive values of the outcome fulfil medical requirements. Another aspect of transferability, when clinical laboratory results are used, concerns differences in analytical performance of measuring instruments. The relative bias between two instruments is not known exactly, but must be estimated and corrected for. A general method, based on original measured data sets and statistical modeling, was developed for simulating the impact of various correction procedures when using different analytical instruments. The simulation methodology was applied to a real clinical problem of ruling-in/ruling-out of patients with suspected acute myocardial infarction (AMI) by biochemical monitoring. The recommended correction procedure was based on method comparison with use of five duplicate measurements on a common set of patient samples covering the relevant measuring interval. Transferability of laboratory data over time was also studied. The design of quality assurance procedures should be based on analytical quality requirement specifications related to medical needs. Limits of critically sized systematic errors were assessed by calculating the decrease in diagnostic performance of the ANN-algorithm as a result of temporary analytical disturbances. The consequences for the design of QA procedures was illustrated. It is concluded that the actual ANN-decision algorithm for early assessment of chest-pain patients should be possible to transfer to new sites under realistic conditions.

Algorithms↗

Variability of reported headache symptoms and diagnosis of migraine at 12 months.

Assignment of a diagnosis of migraine has been formalized in diagnostic criteria proposed by the International Headache Society. The objective of the present study is to determine the reproductibility of the formal diagnosis of migraine in a cohort of headache sufferers over a one-year period. The study was performed in a community cohort taking part in a long-term prospective health survey, the GAZEL study. Two thousand five hundred individuals reporting headache in the GAZEL cohort were sent two postal questionnaires concerning headache symptoms and features at 12-monthly intervals. Replies to the questions allowed a migraine diagnosis to be attributed retrospectively using an algorithm based on the IHS classification scheme. The response rate was 82% for the first questionnaire and 69% for both questionnaires. Of the 1733 subjects providing information at both time-points, the agreement rate for the diagnosis of strict migraine (IHS categories 1.1 or 1.2) was 77.7% (kappa = 0.48), with 62.2% of the patients with this diagnosis (IHS categories 1.1 or 1.2) at Month 0 retaining the same diagnosis at Month 12. When diagnostic criteria were widened to include IHS category 1.7 (migrainous disorder), the agreement rate of the diagnosis was similar at 77.6% (kappa = 0.52), but 82% of the patients with this diagnosis (IHS categories 1.1 or 1.2 or 1.7) at Month 0 now retained the same diagnosis at Month 12. In conclusion, the one-year reproducibility of reporting of migraine headache symptoms is only moderate, varies between symptoms, and leads to instability in the formal assignment of a migraine headache diagnosis and to diagnostic drift between headache types. This finding is compatible with the continuum model of headache, where headache attacks can vary along a severity continuum from episodic tension-type headaches to full-blown migraine attacks.

Algorithms↗

Gene selection for sample classification based on gene expression data: study of sensitivity to choice of parameters of the GA/KNN method.

MOTIVATION: We recently introduced a multivariate approach that selects a subset of predictive genes jointly for sample classification based on expression data. We tested the algorithm on colon and leukemia data sets. As an extension to our earlier work, we systematically examine the sensitivity, reproducibility and stability of gene selection/sample classification to the choice of parameters of the algorithm. METHODS: Our approach combines a Genetic Algorithm (GA) and the k-Nearest Neighbor (KNN) method to identify genes that can jointly discriminate between different classes of samples (e.g. normal versus tumor). The GA/KNN method is a stochastic supervised pattern recognition method. The genes identified are subsequently used to classify independent test set samples. RESULTS: The GA/KNN method is capable of selecting a subset of predictive genes from a large noisy data set for sample classification. It is a multivariate approach that can capture the correlated structure in the data. We find that for a given data set gene selection is highly repeatable in independent runs using the GA/KNN method. In general, however, gene selection may be less robust than classification. AVAILABILITY: The method is available at http://dir.niehs.nih.gov/microarray/datamining CONTACT: LI3@niehs.nih.gov

Algorithms↗

Automatic atrial tachyarrhythmia detection from intracardiac electrograms.

BACKGROUND: Automatic atrial tachyarrhythmia recognition is crucial in order to allow a correct switching-mode function of dual-chamber pacemakers and to avoid inappropriate shocks of ventricular implantable cardioverter-defibrillators. In this paper we considered three algorithms suitable for implantable devices. The first was based on the atrial cycle length; the others analyze different morphologic characteristics of atrial signals. METHODS: Intracardiac bipolar electrogram recordings were obtained from the high right atrium during electrophysiological study. Twenty patients were considered, some of them presenting with different types of cardiac rhythm at different intervals of the study. Cardiac rhythms were divided into three groups: sinus rhythm consisting of 2,196 s obtained from 12 subjects, atrial fibrillation consisting of 771 s obtained from 7 subjects, and atrial flutter consisting of 1,793 s obtained from 7 subjects. The automatic detection was performed on each electrogram segment lasting 1 or 4 s. Atrial segments were separated into two subgroups: the first for the training of the algorithm and the second for testing and validation of results. We considered two types of statistical analysis: comparison between pairs of rhythm (paired classification), and classification among the three different groups (direct classification). RESULTS: The combination of the cycle length algorithm with a morphological method achieved the best performance for both statistical analyses. Paired classification resulted in the following: atrial fibrillation vs sinus rhythm was detected with no error; atrial flutter vs sinus rhythm with a total accuracy of 99.3% (sensitivity 99.4%, specificity 99.2%); atrial fibrillation vs atrial flutter with a total accuracy of 99.1% (sensitivity 98.5%, specificity 99.4%). The total accuracy achieved for the direct classification was 98.6% (average sensitivity 98.5%, specificity 98.8%). CONCLUSIONS: Our results support the association of algorithms for future enhancement of atrial tachyarrhythmia detection in dual-chamber devices, thanks to the limited computational effort.

Algorithms↗

Interactions among factors affecting stillbirths in Holstein cattle in the United States.

Each year about 7% of the Holstein calves born in the United States die within 48 h of birth. The exact cause of death is unknown. The purpose of this article is to examine the complex interactions among factors (e.g., parity, season of birth, dystocia, year) contributing to stillbirth rates. A modified chi-squared automated interaction detection algorithm was used to develop classification trees explaining the most likely sequence of factors that result in a stillborn calf. The data were 666,341 births from the MidStates Dairy Records Processing Center and the National Association of Animal Breeders. Primiparous and multiparous cows clearly differ in the rate of stillbirths, 11.0 and 5.7%, respectively. Dystocia followed parity as the next most important factor within both primiparous and multiparous cows. In primiparous cows, season, year of birth, or gestation length ranked third as an important predictor for dystocia equal to 1, 2, or 3+, respectively. Gestation length ranked third in importance among the factors that affect stillbirth rates for all levels of dystocia in multiparous cows. Among multiparous cows needing assistance (dystocia 3+), stillbirth rates were greatest for shorter gestations less than the average of 280 d, 55.3% for -15 to -12 d, 45.5% for -11 to -9 d, 33.7% for -8 to -5 d, 23.8% for -4 to 13 d, and 35.4% for 14 to 15 d. Gestation length pinpointed the time when stillbirths occurred, as indicated by the increase from 23.8% stillbirth rate among calves born at or above the mean gestation length to 55.3% for those calves born -15 to -12 d below the mean gestation. Further investigation of the relationship between stillbirth rates and gestation length is needed to develop a more complete understanding of the biological processes resulting in the loss of calves at birth.

Algorithms↗

Classification tree models for the prediction of blood-brain barrier passage of drugs.

The use of classification trees for modeling and predicting the passage of molecules through the blood-brain barrier was evaluated. The models were built and evaluated using a data set of 147 molecules extracted from the literature. In the first step, single classification trees were built and evaluated for their predictive abilities. In the second step, attempts were made to improve the predictive abilities using a set of 150 classification trees in a boosting approach. Two boosting algorithms, discrete and real adaptive boosting, were used and compared. High-predictive classification trees were obtained for the data set used, and the models could be improved with boosting. In the context of this research, discrete adaptive boosting gives slightly better results than real adaptive boosting.

Blood-Brain Barrier↗