PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

Support vector machine implementations for classification & clustering.

BACKGROUND: We describe Support Vector Machine (SVM) applications to classification and clustering of channel current data. SVMs are variational-calculus based methods that are constrained to have structural risk minimization (SRM), i.e., they provide noise tolerant solutions for pattern recognition. The SVM approach encapsulates a significant amount of model-fitting information in the choice of its kernel. In work thus far, novel, information-theoretic, kernels have been successfully employed for notably better performance over standard kernels. Currently there are two approaches for implementing multiclass SVMs. One is called external multi-class that arranges several binary classifiers as a decision tree such that they perform a single-class decision making function, with each leaf corresponding to a unique class. The second approach, namely internal-multiclass, involves solving a single optimization problem corresponding to the entire data set (with multiple hyperplanes). RESULTS: Each SVM approach encapsulates a significant amount of model-fitting information in its choice of kernel. In work thus far, novel, information-theoretic, kernels were successfully employed for notably better performance over standard kernels. Two SVM approaches to multiclass discrimination are described: (1) internal multiclass (with a single optimization), and (2) external multiclass (using an optimized decision tree). We describe benefits of the internal-SVM approach, along with further refinements to the internal-multiclass SVM algorithms that offer significant improvement in training time without sacrificing accuracy. In situations where the data isn't clearly separable, making for poor discrimination, signal clustering is used to provide robust and useful information--to this end, novel, SVM-based clustering methods are also described. As with the classification, there are Internal and External SVM Clustering algorithms, both of which are briefly described.

Algorithms↗

Automatic volumetric measurement of lateral ventricles on magnetic resonance images with correction of partial volume effects.

PURPOSE: To propose a method for the quantification of lateral ventricle (LV) volumes on a single sequence of 3D magnetic resonance (MR) images. MATERIALS AND METHODS: This algorithm, following a preliminary fuzzy tissue classification step, is based on the development of mathematical morphology processes allowing both the extraction of the LVs and the correction of partial volume effects on their boundaries. The procedure is fast and totally unsupervised. The method is tested on a phantom image, then applied to five patients diagnosed as potentially suffering from Alzheimer's disease, and finally applied on several MR acquisitions to show the genericness of the algorithm. RESULTS AND CONCLUSION: This technique yielded both an accurate estimation of ventricular volumes intra- and intersubject with respect to published data and a relevant management of partial volume effects. Numerous clinical applications are now expected, from the study of schizophrenia to the longitudinal follow-up of Alzheimer's patients.

Aged↗

Magnetic resonance classification of lumbar intervertebral disc degeneration.

STUDY DESIGN: A reliability study was conducted. OBJECTIVES: To develop a classification system for lumbar disc degeneration based on routine magnetic resonance imaging, to investigate the applicability of a simple algorithm, and to assess the reliability of this classification system. SUMMARY OF BACKGROUND DATA: A standardized nomenclature in the assessment of disc abnormalities is a prerequisite for a comparison of data from different investigations. The reliability of the assessment has a crucial influence on the validity of the data. Grading systems of disc degeneration based on state of the art magnetic resonance imaging and corresponding reproducibility studies currently are sparse. METHODS: A grading system for lumbar disc degeneration was developed on the basis of the literature. An algorithm to assess the grading was developed and optimized by reviewing lumbar magnetic resonance examinations. The reliability of the algorithm in depicting intervertebral disc alterations was tested on the magnetic resonance images of 300 lumbar intervertebral discs in 60 patients (33 men and 27 women) with a mean age of 40 years (range, 10-83 years). All scans were analyzed independently by three observers. Intra- and interobserver reliabilities were assessed by calculating kappa statistics. RESULTS: There were 14 Grade I, 82 Grade II, 72 Grade III, 68 Grade IV, and 64 Grade V discs. The kappa coefficients for intra- and interobserver agreement were substantial to excellent: intraobserver (kappa range, 0.84-0.90) and interobserver (kappa range, 0.69-0.81). Complete agreement was obtained, on the average, in 83.8% of all the discs. A difference of one grade occurred in 15.9% and a difference of two or more grades in 1.3% of all the cases. CONCLUSION: Disc degeneration can be graded reliably on routine T2-weighted magnetic resonance images using the grading system and algorithm presented in this investigation.

Algorithms↗

Methodology of ECG interpretation in the Hannover program.

The Hannover ECG program HES has been designed for measurement and interpretation of resting and (moderate) exercise electrocardiograms. In the signal analysis part the program follows an averaging strategy. For diagnostic classification a hybrid model with decision trees and scoring algorithms, and with multivariate probabilistic tests for derivation of category A statements is applied. The multivariate classification technique allows to adjust sensitivity and specificity for specific application areas without changing the diagnostic criteria.

Algorithms↗

Nonparametric supervised learning by linear interpolation with maximum entropy.

Nonparametric neighborhood methods for learning entail estimation of class conditional probabilities based on relative frequencies of samples that are "near-neighbors" of a test point. We propose and explore the behavior of a learning algorithm that uses linear interpolation and the principle of maximum entropy (LIME). We consider some theoretical properties of the LIME algorithm: LIME weights have exponential form; the estimates are consistent; and the estimates are robust to additive noise. In relation to bias reduction, we show that near-neighbors contain a test point in their convex hull asymptotically. The common linear interpolation solution used for regression on grids or look-up-tables is shown to solve a related maximum entropy problem. LIME simulation results support use of the method, and performance on a pipeline integrity classification problem demonstrates that the proposed algorithm has practical value.

Algorithms↗

Peak pressure curve: an effective parameter for early detection of foot functional impairments in diabetic patients.

A clinical investigation was conducted on 61 diabetic patients and 22 healthy volunteers. Joint mobility, muscular function of the foot-ankle complex and plantar pressure measurements were characterised. A clustering algorithm was applied to obtain patient classification based on the shape and amplitude of the time curve of the instantaneous maximum pressure the foot experienced during gait. Results indicate that a screening test based on the peak pressure curve might be an effective way to detect diabetic patients at risk of foot ulceration.

Algorithms↗

Probabilistic inference-based classification applied to myoelectric signal decomposition.

A new probabilistic inference based technique (IBC) for the classification of motor unit action potentials (MUAP's) is presented. This new technique discovers statistically significant relationships in the data and uses these relationships to generate classification rules. The technique was applied to the classification of MUAP's extracted from simulated myoelectric signals. Its performance was compared to that of classical template matching algorithms (TBC) applied to the same data. Using 32 time samples as features to represent the MUAP's it was found that the IBC based technique performed significantly better (p less than 0.005) than the TBC algorithms (83.0 +/- 2.6% versus 78.1 +/- 2.8% peak correct classification performance). As the size of the training set was reduced or as increasing numbers of random classification errors were introduced into the training data, the performance of the IBC and TBC techniques declined similarly. IBC performance remained superior until very small training sets (less than 30 MUAP's per motor unit) or training sets with large numbers of errors (greater than 50%) were used. Because the probabilistic inference technique can utilize nominal data it has the potential to use declarative problem domain knowledge which conceivably could improve its performance.

Action Potentials↗

Molecular descriptors for effective classification of biologically active compounds based on principal component analysis identified by a genetic algorithm.

We have evaluated combinations of 111 descriptors that were calculated from two-dimensional representations of molecules to classify 455 compounds belonging to seven biological activity classes using a method based on principal component analysis. The analysis was facilitated by application of a genetic algorithm. Using scoring functions that related the number of compounds in pure classes (i.e., compounds with the same biological activity), singletons, and mixed classes, effective descriptor sets were identified. A combination of only four molecular descriptors accounting for aromatic character, hydrogen bond acceptors, estimated polar van der Waals surface area, and a single structural key gave overall best results. At this performance level, approximately 91% of the compounds occurred in pure classes and mixed classes were absent. The results indicate that combinations of only a few critical descriptors are preferred to partition compounds according to their biological activity, at least in the test cases studied here.

Algorithms↗

Study of sleep-wakefulness states by computer graphics and cluster analysis before and after lesions of the pontine tegmentum in the cat.

A computerized method of quantification, graphic representation and classification of sleep-wakefulness data in the cat before and after pontine tegmental lesions has been presented. Electrophysiological signal features including average EEG amplitude, average EMG amplitude and PGO spike rate which are particularly important for the definition of sleep-wakefulness states have been quantified for each 1 min epoch in the day. The data were presented in a projected 3-dimensional data display, in which they formed clusters that are considered to be analogous to sleep-wakefulness states. A cluster analysis algorithm was employed for the automatic classification of these data, and this automatic classification was compared graphically and with contingency table analyses to traditional visual assessment of state from polygraphic records. Although there were systematic differences in the locations of state boundaries, total percent agreement between cluster analysis classification and traditional human classification was comparable to the percent agreement between any two human classifiers (about 90%). After pontine tegmental lesions involving both the gigantocellular and lateral tegmental fields, paradoxical sleep was eliminated, and the characteristics of slow wave sleep and wakefulness were altered. The elimination of the state of paradoxical sleep was evident in the computer display by the absence of the paradoxical sleep cluster, and alterations of the other states were indicated in the display by shifts in the positions of their respective clusters. Automatic classification of slow wave sleep and wakefulness after such lesions compared well with traditional classification, attesting to the validity of this approach.

Animals↗

Management and return to play of stress fractures.

OBJECTIVE: The purpose of this article is to provide the clinician an evidence/experience-based algorithm for the management of stress fractures. DATA SOURCES: Medline search of peer reviewed publications regarding stress fracture etiology, classification, treatment, and natural history. DATA SYNTHESIS/METHODS: The algorithm was developed from a review of retrospective case series, a few evidence-based papers, and the clinical experience of 4 sports medicine team physicians with a combined experience of over 40 years in the care of athletes at the college and professional level. The literature is almost entirely case series without control groups; therefore, clinical consensus is included as the next best guide to treatment. RESULTS: The emphasis of this article is to provide a clear and simple approach to the management of these fractures by classifying them as either high-risk or low-risk. This separation into 2 groups is based on the biomechanical environment and natural history of the fracture. High-risk stress fractures occur in the superolateral femoral neck, anterior tibial shaft, tarsal navicular, proximal fifth metatarsal, and talar neck. Low-risk stress fractures occur in the lateral malleolus, calcaneus, 2nd through 4th metatarsals, and the femoral shaft. CONCLUSIONS: The undertreatment of high-risk stress fractures can lead to catastrophic bone failure and/or prolonged loss of playing time. Overtreatment of low-risk stress fractures can result in unnecessary deconditioning and unneeded loss of playing time. We propose that the use of the simple and clinically relevant algorithm will help guide appropriate management and return to play decision-making as well as encourage future prospective research.

Algorithms↗

A selective mapping algorithm for computer analysis of voided urine cell images.

One of the fundamental targets of the automated image analysis of cytologic preparations is the reduction of computer classification errors due to cells or other objects that do not lend themselves to image segmentation or that have morphologic features that may mislead the cell classification schemes. In prior work from this laboratory, the achievement of this goal was attempted by hierarchical analysis of sequential microscopic objects at high resolution. This paper reports on the successful development and implementation of an automated "selective mapping algorithm" that selects cells at low power for further analysis and eliminates a large proportion of unwanted "objects." The algorithm classifies the objects and extracts appropriate features from a 256 X 240 digital image obtained via a 10 X planachromatic objective. The five-node binary tree classifier used in this triage is described. The algorithm was trained and tested initially on 501 visually classified microscopic "objects," resulting in a correct acceptance rate of 61.3% and correct rejection rate of 81.3%. The selective mapping algorithm was subsequently integrated into the video-based image analysis system constructed at the Montefiore Medical Center for the diagnostic evaluation of sediments of voided urine. The algorithm was then tested on ten cytocentrifuge preparations for a preliminary evaluation of its performance. Up to 100 "objects" per case were selected by the algorithm for further classification by the computer at high power. Of the 810 "objects" selected by the selective mapping algorithm, 344 (42.5%) were classified by the computer at high resolution as cells of diagnostic value ("WELL" cells) and 466 were rejected.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗

An entropy-based gene selection method for cancer classification using microarray data.

BACKGROUND: Accurate diagnosis of cancer subtypes remains a challenging problem. Building classifiers based on gene expression data is a promising approach; yet the selection of non-redundant but relevant genes is difficult. The selected gene set should be small enough to allow diagnosis even in regular clinical laboratories and ideally identify genes involved in cancer-specific regulatory pathways. Here an entropy-based method is proposed that selects genes related to the different cancer classes while at the same time reducing the redundancy among the genes. RESULTS: The present study identifies a subset of features by maximizing the relevance and minimizing the redundancy of the selected genes. A merit called normalized mutual information is employed to measure the relevance and the redundancy of the genes. In order to find a more representative subset of features, an iterative procedure is adopted that incorporates an initial clustering followed by data partitioning and the application of the algorithm to each of the partitions. A leave-one-out approach then selects the most commonly selected genes across all the different runs and the gene selection algorithm is applied again to pare down the list of selected genes until a minimal subset is obtained that gives a satisfactory accuracy of classification. The algorithm was applied to three different data sets and the results obtained were compared to work done by others using the same data sets. CONCLUSION: This study presents an entropy-based iterative algorithm for selecting genes from microarray data that are able to classify various cancer sub-types with high accuracy. In addition, the feature set obtained is very compact, that is, the redundancy between genes is reduced to a large extent. This implies that classifiers can be built with a smaller subset of genes.

Algorithms↗

[Bronchoscopic assessment algorithms for the practical evaluation of the rheological properties of the tracheobronchial secretion and the classification of the degree of the disordered drainage function of the tracheobronchial tree (TBT) in chest and combined trauma with chest trauma as the leading injury].

Ventilation impairment, due to ineffective elimination of the mucous-hemorrhagic content from the tracheobronchial tree (TBT), obstructs the upper airways with the ensuing ventilation reduction giving rise to atelectases and progressive alveolar block. There is evidence of transudation and exudation into the pulmonary pathways and pleural cavity. A series of 276 patients presenting closed chest trauma are subjected to fibrobronchoscopy (FBS) and follow-up study. In 92 of them bronchoscopy is performed 2 to 15 times per patient, accordingly: in 75-twice, in 10-five times and in 15-twice. One-hundred twenty-nine of the total of 276 cases under study are on mechanical ventilation. In 56 instances FBS is carried out through a tracheostomy cannula, in 73-by intubation, in 18-through the mouth, and in two--through the nose. Based on the obtained results, algorithms for assessment of the rheological properties of tracheobronchial secretion and degree of impairment of TBT drainage function during emergency FBS in closed chest injuries are worked out, having an essential practical bearing on the diagnostic and therapeutic approach to closed thoracic trauma.

Algorithms↗

Genes for the majority of group a streptococcal virulence factors and extracellular surface proteins do not confer an increased propensity to cause invasive disease.

BACKGROUND: The factors behind the reemergence of severe, invasive group A streptococcal (GAS) diseases are unclear, but it could be caused by altered genetic endowment in these organisms. However, data from previous studies assessing the association between single genetic factors and invasive disease are often conflicting, suggesting that other, as-yet unidentified factors are necessary for the development of this class of disease. METHODS: In this study, we used a targeted GAS virulence microarray containing 226 GAS genes to determine the virulence gene repertoires of 68 GAS isolates (42 associated with invasive disease and 28 associated with noninvasive disease) collected in a defined geographic location during a contiguous time period. We then employed 3 advanced machine learning methods (genetic algorithm neural network, support vector machines, and classification trees) to identify genes with an increased association with invasive disease. RESULTS: Virulence gene profiles of individual GAS isolates varied extensively among these geographically and temporally related strains. Using genetic algorithm neural network analysis, we identified 3 genes with a marginal overrepresentation in invasive disease isolates. Significantly, 2 of these genes, ssa and mf4, encoded superantigens but were only present in a restricted set of GAS M-types. The third gene, spa, was found in variable distributions in all M-types in the study. CONCLUSIONS: Our comprehensive analysis of GAS virulence profiles provides strong evidence for the incongruent relationships among any of the 226 genes represented on the array and the overall propensity of GAS to cause invasive disease, underscoring the pathogenic complexity of these diseases, as well as the importance of multiple bacteria and/or host factors.

Humans↗

Classification of faces in man and machine.

We attempt to shed light on the algorithms humans use to classify images of human faces according to their gender. For this, a novel methodology combining human psychophysics and machine learning is introduced. We proceed as follows. First, we apply principal component analysis (PCA) on the pixel information of the face stimuli. We then obtain a data set composed of these PCA eigenvectors combined with the subjects' gender estimates of the corresponding stimuli. Second, we model the gender classification process on this data set using a separating hyperplane (SH) between both classes. This SH is computed using algorithms from machine learning: the support vector machine (SVM), the relevance vector machine, the prototype classifier, and the K-means classifier. The classification behavior of humans and machines is then analyzed in three steps. First, the classification errors of humans and machines are compared for the various classifiers, and we also assess how well machines can recreate the subjects' internal decision boundary by studying the training errors of the machines. Second, we study the correlations between the rank-order of the subjects' responses to each stimulus-the gender estimate with its reaction time and confidence rating-and the rank-order of the distance of these stimuli to the SH. Finally, we attempt to compare the metric of the representations used by humans and machines for classification by relating the subjects' gender estimate of each stimulus and the distance of this stimulus to the SH. While we show that the classification error alone is not a sufficient selection criterion between the different algorithms humans might use to classify face stimuli, the distance of these stimuli to the SH is shown to capture essentials of the internal decision space of humans. Furthermore, algorithms such as the prototype classifier using stimuli in the center of the classes are shown to be less adapted to model human classification behavior than algorithms such as the SVM based on stimuli close to the boundary between the classes.

Algorithms↗

A fast and accurate online sequential learning algorithm for feedforward networks.

In this paper, we develop an online sequential learning algorithm for single hidden layer feedforward networks (SLFNs) with additive or radial basis function (RBF) hidden nodes in a unified framework. The algorithm is referred to as online sequential extreme learning machine (OS-ELM) and can learn data one-by-one or chunk-by-chunk (a block of data) with fixed or varying chunk size. The activation functions for additive nodes in OS-ELM can be any bounded nonconstant piecewise continuous functions and the activation functions for RBF nodes can be any integrable piecewise continuous functions. In OS-ELM, the parameters of hidden nodes (the input weights and biases of additive nodes or the centers and impact factors of RBF nodes) are randomly selected and the output weights are analytically determined based on the sequentially arriving data. The algorithm uses the ideas of ELM of Huang et al. developed for batch learning which has been shown to be extremely fast with generalization performance better than other batch training methods. Apart from selecting the number of hidden nodes, no other control parameters have to be manually chosen. Detailed performance comparison of OS-ELM is done with other popular sequential learning algorithms on benchmark problems drawn from the regression, classification and time series prediction areas. The results show that the OS-ELM is faster than the other sequential algorithms and produces better generalization performance.

Algorithms↗

From image processing to classification: II. Classification of electrophoretic patterns using self-organizing feature maps and feed-forward neural networks.

In a recent study, isoelectric focusing patterns were classified with a neural network using the back-propagation algorithm [1]. In order to further study the classification process and to generalize the presentation of electrophoretic patterns, Kohonen's self-organizing feature maps [2] were applied in this study. Although these feature maps are very efficient in many pattern recognition tasks, our data proved to be too complex for classification with an unsupervised system. Therefore, a second supervised network on top of the feature map was necessary. As in [3], a feed-forward network trained by the back-propagation algorithm was used. The final system allows us to correctly classify 90% of all wheat varieties. Moreover, the system proved to be reliable, reasonable in training time and shows the same accuracy in different experimental setups.

Algorithms↗

Stabilized binary hierarchic classifier in cytopathologic diagnosis.

A binary tree classifier (BTC) algorithm for computer-assisted cell image analysis has been developed that overcomes the problem of overtraining due to inadequate sample size/dimensionality ratio at the higher-order nodes of a hierarchic decision structure. Provisions have been introduced that ensure that decision rules created at each node are based on samples representative of the subpopulation routed to the node. These provisions eliminate problems caused by truncation effects resulting from the application of decision rules at preceding decision nodes. The BTC performs better than do single-stage classifiers in situations where the categories' mean vectors are not well separated and no equality of covariance matrices exists. In applications in which noticeable deterioration of classifier performance on test-set data is common, the classification success rate of the BTC algorithm is not statistically significantly different between the training-set and test-set data.

Cells↗