PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

A new algorithm for the evaluation of shotgun peptide sequencing in proteomics: support vector machine classification of peptide MS/MS spectra and SEQUEST scores.

Shotgun tandem mass spectrometry-based peptide sequencing using programs such as SEQUEST allows high-throughput identification of peptides, which in turn allows the identification of corresponding proteins. We have applied a machine learning algorithm, called the support vector machine, to discriminate between correctly and incorrectly identified peptides using SEQUEST output. Each peptide was characterized by SEQUEST-calculated features such as delta Cn and Xcorr, measurements such as precursor ion current and mass, and additional calculated parameters such as the fraction of matched MS/MS peaks. The trained SVM classifier performed significantly better than previous cutoff-based methods at separating positive from negative peptides. Positive and negative peptides were more readily distinguished in training set data acquired on a QTOF, compared to an ion trap mass spectrometer. The use of 13 features, including four new parameters, significantly improved the separation between positive and negative peptides. Use of the support vector machine and these additional parameters resulted in a more accurate interpretation of peptide MS/MS spectra and is an important step toward automated interpretation of peptide tandem mass spectrometry data in proteomics.

Algorithms↗

Boosted mixture of experts: an ensemble learning scheme.

We present a new supervised learning procedure for ensemble machines, in which outputs of predictors, trained on different distributions, are combined by a dynamic classifier combination model. This procedure may be viewed as either a version of mixture of experts (Jacobs, Jordan, Nowlan, & Hintnon, 1991), applied to classification, or a variant of the boosting algorithm (Schapire, 1990). As a variant of the mixture of experts, it can be made appropriate for general classification and regression problems by initializing the partition of the data set to different experts in a boostlike manner. If viewed as a variant of the boosting algorithm, its main gain is the use of a dynamic combination model for the outputs of the networks. Results are demonstrated on a synthetic example and a digit recognition task from the NIST database and compared with classifical ensemble approaches.

Algorithms↗

A geometric approach to support vector machine (SVM) classification.

The geometric framework for the support vector machine (SVM) classification problem provides an intuitive ground for the understanding and the application of geometric optimization algorithms, leading to practical solutions of real world classification problems. In this work, the notion of "reduced convex hull" is employed and supported by a set of new theoretical results. These results allow existing geometric algorithms to be directly and practically applied to solve not only separable, but also nonseparable classification problems both accurately and efficiently. As a practical application of the new theoretical results, a known geometric algorithm has been employed and transformed accordingly to solve nonseparable problems successfully.

Algorithms↗

Classification of normal and dysphagic swallows by acoustical means.

This paper proposes a noninvasive, acoustic-based method to differentiate between individuals with and without dysphagia or swallowing dysfunction. Swallowing sound signals, both normal and abnormal (i.e., at risk of some degree of dysphagia) were recorded with accelerometers over the trachea. Segmentation based on waveform dimension trajectory (a distance-based technique) was developed to segment the nonstationary swallowing sound signals. Two characteristic sections emerged, Opening and Transmission, and 24 characteristic features were extracted and subsequently reduced via discriminant analysis. A discriminant algorithm was also employed for classification, with the system trained and tested using the leave-one-out approach. Overall, 350 signals were used from three bolus consistencies (semisolid, thick and thin liquids). A final screening algorithm correctly classified 13 of 15 control subjects and 11 of 11 subjects with some degree of dysphagia and/or neurological impairments. The proposed method has great potential to reduce the need for videofluoroscopic swallowing studies (the current gold standard method for swallowing assessment, which is invasive and nonportable) and to assist in the overall clinical assessment of swallowing sound signals.

Adolescent↗

An artificial intelligent diagnostic system on differential recognition of hematopoietic cells from microscopic images.

Despite their advantages, none of the automated white blood cell differentiated counters have replaced the conventional microscopic evaluations of blood and bone marrow slides by hematologists. We have analyzed the smears of 39 patients and 8 control subjects to develop an artificial expert system that recognizes 16 different types of nucleated hematopoietic cells during the stages of differentiation. A charge coupled television camera and a special frame grabber were used for data acquisition, and 247 nucleated cell images were transferred from a microscope to an IBM 386 computer to be processed. One hundred sixty-five and 82 of these images were used for training and testing, respectively. Our system is composed of image processing and analysis (enhancement, thresholding/smoothing, edge detection), pattern recognition (feature extraction and classification with supervised artificial neural network), and expert system development. Image processing and analysis were used to obtain 13 cellular features to be used as the input parameters (neurons) of the artificial neural network. A supervised artificial neural network (back-propagation learning algorithm) was used in the classification of 16 different cells (output neurons of the neural network), which is the second step of pattern recognition. A confusion matrix has been developed to compare the similarities and dissimilarities between the differential recognitions of the hematologist and the expert system. The discriminatory power of the procedure is statistically significant: Q = (N - n.K)2/N.(K - 1) = 28.2. The sensitivity and the specificity of the expert system were 71.4% and 90.9%, respectively.

Bone Marrow↗

Cluster analysis by testing the statistical hypothesis of uniformity.

This paper presents a cluster algorithm that defines the number of clusters and allows classification of data points. The basic task of the algorithm is to identify accumulations of vectors in the analysis sample of vectors. The accumulations of vectors are determined by testing the statistical hypothesis of uniformity. On the basis of accumulations, clusters are formed.

Algorithms↗

Sensitivity and specificity of confocal laser-scanning microscopy for in vivo diagnosis of malignant skin tumors.

BACKGROUND: Melanoma and nonmelanoma skin cancer are the most frequent malignant tumors by far among whites. Currently, early diagnosis is the most efficient method for preventing a fatal outcome. In vivo confocal laser-scanning microscopy (CLSM) is a recently developed potential diagnostic tool. METHODS: One hundred seventeen melanocytic skin lesions and 45 nonmelanocytic skin lesions (90 benign nevi, 27 malignant melanomas, 15 basal cell carcinomas, and 30 seborrheic keratoses) were sampled consecutively and were examined using proprietary CLSM equipment. Stored images were rated by 4 independent observers. RESULTS: Differentiation between melanoma and all other lesions based solely on CLSM examination was achieved with a positive predictive value of 94.22%. Malignant lesions (melanoma and basal cell carcinoma) as a group were diagnosed with a positive predictive value of 96.34%. Assessment of distinct CLSM features showed a strong interobserver correlation (kappa >0.80 for 11 of 13 criteria). Classification and regression tree analysis yielded a 3-step algorithm based on only 3 criteria, facilitating a correct classification in 96.30% of melanomas, 98.89% of benign nevi, and 100% of basal cell carcinomas and seborrheic keratoses. CONCLUSIONS: In vivo CLSM examination appeared to be a promising method for the noninvasive assessment of melanoma and nonmelanoma skin tumors.

Basal Cell Carcinoma↗

Short fuzzy tandem repeats in genomic sequences, identification, and possible role in regulation of gene expression.

MOTIVATION: Genomic sequences are highly redundant and contain many types of repetitive DNA. Fuzzy tandem repeats (FTRs) are of particular interest. They are found in regulatory regions of eukaryotic genes and are reported to interact with transcription factors. However, accurate assessment of FTR occurrences in different genome segments requires specific algorithm for efficient FTR identification and classification. RESULTS: We have obtained formulas for P-values of FTR occurrence and developed an FTR identification algorithm implemented in TandemSWAN software. Using TandemSWAN we compared the structure and the occurrence of FTRs with short period length (up to 24 bp) in coding and non-coding regions including UTRs, heterochromatic, intergenic and enhancer sequences of Drosophila melanogaster and Drosophila pseudoobscura. Tandems with period three and its multiples were found in coding segments, whereas FTRs with periods multiple of six are overrepresented in all non-coding segment. Periods equal to 5-7 and 11-14 were characteristic of the enhancer regions and other non-coding regions close to genes. AVAILABILITY: TandemSWAN web page, stand-alone version and documentation can be found at http://bioinform.genetika.ru/projects/swan/www/ SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

Clustering algorithms and other exploratory methods for microarray data analysis.

OBJECTIVES: We introduce methods for the exploratory analysis of microarray data, especially focusing on cluster algorithms. Benefits and problems are discussed. METHODS: We describe application and suitability of unsupervised learning methods for the classification of gene expression data. Cluster algorithms are treated in more detail, including assessment of cluster quality. RESULTS: When dealing with microarray data, most cluster algorithms must be applied with caution. As long as the structure of the true generating models of such data is not fully understood, the use of simple algorithms seems to be more appropriate than the application of complex black-box algorithms. New methods explicitly targeted to the analysis of microarray data are increasingly being developed in order to increase the amount of useful information extracted from the experiments. CONCLUSIONS: Unsupervised methods can be a helpful tool for the analysis of microarray data, but a critical choice of the algorithm and a careful interpretation of the results are required in order to avoid false conclusions.

Algorithms↗

Likelihood linkage analysis (LLA) classification method: an example treated by hand.

This paper describes a very general method of data analysis using a hierarchical classification. The data can be provided by observation, experiment or knowledge; their nature can be numerical, qualitative or logical. First, the classical view of the context of data representation, in which the algorithm of hierarchical ascendant construction of the classification tree is set, is treated in a synthetic manner. The main notion in our method is one of 'similarity'. This must be elaborated in the best way, taking into account the mathematical nature of the objects to be compared. Here we adopt a set of theoretical and combinatorial representation of the descriptive attributes, which are interpreted in terms of relations. Then we introduce a probability scale for similarity measurement by using a likelihood concept. The largest part of the paper concerns an illustrating example, moderately sized, detailing very minutely the different steps and the different calculations assumed by the method. The data structure handled with this example is the simplest possible. Then, general aspects and methodological extensions are evoked. We end by indicating the interest of the described approach in future works, in which we are involved, concerning typological organization of genetic sequences. We emphasize the 'explanation' aspect of the obtained results, with respect to a given description. For this purpose, classifications (on the object set and on the attribute set) on the one hand and machine learning techniques on the other, intervene efficiently.

Algorithms↗

Band features as classification measures for G-banded chromosome analysis.

Modern automatic and semiautomatic karyotyping systems employ algorithms that use chromosome length and centromeric index as well as other intact chromosome measures. These measures offer correct classification rates near 95%. An algorithm is presented that utilizes local dark band features and position (position from one end of the chromosome, band-width, band-height above light band background, integrated optical density above light band background, and a shape feature) and is based on maximum likelihood of the multivariate normal distribution for the feature vector. The algorithm was tested on two data sets: 179 metaphases from C. Lundsteen at the Rigshospitalet, Copenhagen, and 50 metaphases from The University of Texas M. D. Anderson Cancer Center. The Copenhagen set achieved an overall correct classification rate of 94.6% when classifying itself, a rate comparable to other algorithms. This classifier relies on local band features rather than global chromosome characteristics and is therefore directly extensible to metaphase and prophase chromosome subsegments and to structural abnormalities.

Algorithms↗

BCI Competition 2003--Data set III: probabilistic modeling of sensorimotor mu rhythms for classification of imaginary hand movements.

Brain-computer interfaces require effective online processing of electroencephalogram (EEG) measurements, e.g., as a part of feedback systems. We present an algorithm for single-trial online classification of imaginary left and right hand movements, based on time-frequency information derived from filtering EEG wideband raw data with causal Morlet wavelets, which are adapted to individual EEG spectra. Since imaginary hand movements lead to perturbations of the ongoing pericentral mu rhythm, we estimate probabilistic models for amplitude modulation in lower (10 Hz) and upper (20 Hz) frequency bands over the sensorimotor hand cortices both contra- and ipsilaterally to the imagined movements (i.e., at EEG channels C3 and C4). We use an integrative approach to accumulate over time evidence for the subject's unknown motor intention. Disclosure of test data labels after the competition showed this approach to succeed with an error rate as low as 10.7%.

Algorithms↗

Absence of the septum pellucidum: a useful sign in the diagnosis of congenital brain malformations.

In a review of more than 2000 MR images of the brain we identified 35 patients with absence of the septum pellucidum. These patients were divided into seven basic groups as follows: septooptic dysplasia; schizencephaly; holoprosencephaly; agenesis of the corpus callosum; chronic, severe hydrocephalus; basilar encephaloceles; and porencephaly/hydranencephaly. Absence of the septum pellucidum was never seen as an isolated finding. By using data gathered from the review of the MR scans of patients in this study, we devised a diagnostic algorithm to aid in the classification of these patients. Absence of the septum pellucidum can provide a valuable clue to the diagnosis of malformations of the brain.

Abnormalities, Multiple↗

Hierarchical cluster analysis applied to workers' exposures in fiberglass insulation manufacturing.

The objectives of this study were to explore the application of cluster analysis to the characterization of multiple exposures in industrial hygiene practice and to compare exposure groupings based on the result from cluster analysis with that based on non-measurement-based approaches commonly used in epidemiology. Cluster analysis was performed for 37 workers simultaneously exposed to three agents (endotoxin, phenolic compounds and formaldehyde) in fiberglass insulation manufacturing. Different clustering algorithms, including complete-linkage (or farthest-neighbor), single-linkage (or nearest-neighbor), group-average and model-based clustering approaches, were used to construct the tree structures from which clusters can be formed. Differences were observed between the exposure clusters constructed by these different clustering algorithms. When contrasting the exposure classification based on tree structures with that based on non-measurement-based information, the results indicate that the exposure clusters identified from the tree structures had little in common with the classification results from either the traditional exposure zone or the work group classification approach. In terms of the defining homogeneous exposure groups or from the standpoint of health risk, some toxicological normalization in the components of the exposure vector appears to be required in order to form meaningful exposure groupings from cluster analysis. Finally, it remains important to see if the lack of correspondence between exposure groups based on epidemiological classification and measurement data is a peculiarity of the data or a more general problem in multivariate exposure analysis.

Air Pollutants, Occupational↗

A quantitative analysis approach for cardiac arrhythmia classification using higher order spectral techniques.

Ventricular tachyarrhythmias, in particular ventricular fibrillation (VF), are the primary arrhythmic events in the majority of patients suffering from sudden cardiac death. Attention has focused upon these articular rhythms as it is recognized that prompt therapy can lead to a successful outcome. There has been considerable interest in analysis of the surface electrocardiogram (ECG) in VF centred on attempts to understand the pathophysiological processes occurring in sudden cardiac death, predicting the efficacy of therapy, and guiding the use of alternative or adjunct therapies to improve resuscitation success rates. Atrial fibrillation (AF) and ventricular tachycardia (VT) are other types of tachyarrhythmias that constitute a medical challenge. In this paper, a high order spectral analysis technique is suggested for quantitative analysis and classification of cardiac arrhythmias. The algorithm is based upon bispectral analysis techniques. The bispectrum is estimated using an autoregressive model, and the frequency support of the bispectrum is extracted as a quantitative measure to classify atrial and ventricular tachyarrhythmias. Results show a significant difference in the parameter values for different arrhythmias. Moreover, the bicoherency spectrum shows different bicoherency values for normal and tachycardia patients. In particular, the bicoherency indicates that phase coupling decreases as arrhythmia kicks in. The simplicity of the classification parameter and the obtained specificity and sensitivity of the classification scheme reveal the importance of higher order spectral analysis in the classification of life threatening arrhythmias. Further investigations and modification of the classification scheme could inherently improve the results of this technique and predict the instant of arrhythmia change.

Algorithms↗

Automated 3-D extraction and evaluation of the inner and outer cortical surfaces using a Laplacian map and partial volume effect classification.

Accurate reconstruction of the inner and outer cortical surfaces of the human cerebrum is a critical objective for a wide variety of neuroimaging analysis purposes, including visualization, morphometry, and brain mapping. The Anatomic Segmentation using Proximity (ASP) algorithm, previously developed by our group, provides a topology-preserving cortical surface deformation method that has been extensively used for the aforementioned purposes. However, constraints in the algorithm to ensure topology preservation occasionally produce incorrect thickness measurements due to a restriction in the range of allowable distances between the gray and white matter surfaces. This problem is particularly prominent in pediatric brain images with tightly folded gyri. This paper presents a novel method for improving the conventional ASP algorithm by making use of partial volume information through probabilistic classification in order to allow for topology preservation across a less restricted range of cortical thickness values. The new algorithm also corrects the classification of the insular cortex by masking out subcortical tissues. For 70 pediatric brains, validation experiments for the modified algorithm, Constrained Laplacian ASP (CLASP), were performed by three methods: (i) volume matching between surface-masked gray matter (GM) and conventional tissue-classified GM, (ii) surface matching between simulated and CLASP-extracted surfaces, and (iii) repeatability of the surface reconstruction among 16 MRI scans of the same subject. In the volume-based evaluation, the volume enclosed by the CLASP WM and GM surfaces matched the classified GM volume 13% more accurately than using conventional ASP. In the surface-based evaluation, using synthesized thick cortex, the average difference between simulated and extracted surfaces was 4.6 +/- 1.4 mm for conventional ASP and 0.5 +/- 0.4 mm for CLASP. In a repeatability study, CLASP produced a 30% lower RMS error for the GM surface and a 8% lower RMS error for the WM surface compared with ASP.

Algorithms↗

The predictors of pelvic lymph node metastasis at radical retropubic prostatectomy.

PURPOSE: We studied preoperative variables in a contemporary series of patients who underwent radical retropubic prostatectomy (RRP) to determine which variables were associated with lymph node metastasis. MATERIALS AND METHODS: Between January 1995 and November 1999, 1,091 men underwent RRP, 695 of whom underwent bilateral pelvic lymph node dissection without any prior therapy. We evaluated biopsy Gleason score, maximum tumor length and maximum percentage of tumor in the positive core(s), location and number of positive cores, and total prostate specific antigen before surgery in 295 of these patients. We also developed a classification and regression tree analysis algorithm to segregate the risk of positive lymph node metastasis. Stepwise logistic regression analyses were used to determine independent predictors of lymph node metastasis. RESULTS: Of the 695 patients 19 (2.7%) had lymph node metastasis. Clinical stage, Gleason score, positive basal core, greatest percentage of tumor on positive cores and maximum tumor length in positive core were significant predictors of lymph node metastasis in the Mann-Whitney U test and chi-square test. Classification and regression trees analysis revealed that 4 or more positive cores with any Gleason grade 4 or 5, serum prostate specific antigen 15.0 ng/ml or greater, or the presence of dominant Gleason 4 or 5 were independent predictors of lymph node metastasis. Our algorithm had a significantly higher diagnostic performance than the Hamburg algorithm (p = 0.002). CONCLUSIONS: Our algorithm may be a valid tool for the prediction of lymph node metastasis and may help to select men who do not need to undergo bilateral pelvic lymph node dissection with RRP.

Adult↗

Automated severity classification of AIDS hospitalizations.

To validate an automated AIDS severity-of-illness prognostic algorithm, 2,113 discharge summaries of HIV-infected patients were merged with the Problem-Oriented Medical Synopsis (POMS) and an HIV risk registry. The combination of a medically derived classification and staging algorithm with multivariate statistical techniques was used for automated severity-of-illness disease staging and prognostic assignment. The model correctly predicted the outcomes of 82% of all cases (death, survivorship) at discharge, and 66% of deaths.

Algorithms↗