PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Classification Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

An incremental network for on-line unsupervised classification and topology learning.

This paper presents an on-line unsupervised learning mechanism for unlabeled data that are polluted by noise. Using a similarity threshold-based and a local error-based insertion criterion, the system is able to grow incrementally and to accommodate input patterns of on-line non-stationary data distribution. A definition of a utility parameter, the error-radius, allows this system to learn the number of nodes needed to solve a task. The use of a new technique for removing nodes in low probability density regions can separate clusters with low-density overlaps and dynamically eliminate noise in the input data. The design of two-layer neural network enables this system to represent the topological structure of unsupervised on-line data, report the reasonable number of clusters, and give typical prototype patterns of every cluster without prior conditions such as a suitable number of nodes or a good initial codebook.

Algorithms↗

Identifying undiagnosed primary immunodeficiency diseases in minority subjects by using computer sorting of diagnosis codes.

BACKGROUND: Primary immunodeficiency diseases occur in all populations, but these diagnoses are rarely made in minority subjects in the United States. OBJECTIVE: We sought to develop and validate a method to identify patients without diagnoses but with immunodeficiency in an urban hospital with a substantial minority patient population. METHODS: We developed a scoring algorithm on the basis of International Classification of Disease, Ninth Revision (ICD-9) codes to identify all hospitalized patients age 60 years or less who had been given a diagnosis of 2 or more of 174 ICD-9-coded complications associated with immunodeficiency. Codes were weighted for severity and expressed as a sum for all admissions between October 1, 1995, and December 31, 2002. Patients with, for example, cancer or HIV or those after transplantation or major surgery were excluded. Demographic features of subjects with aggregated ICD-9 codes suggestive of immunodeficiency were compared with those of other inpatients; 59 computer-selected subjects were then tested for immune defects. RESULTS: The computer-identified group contained 533 patients (0.4% of all inpatients), who had been hospitalized 2683 times. The median age was 6.6 years. Sixty-five percent were African American or Hispanic, and 61% were insured by Medicaid, which is significantly more than other inpatients younger than 60 years of age (median age, 32.6 years; 37% minority, 27% insured by Medicaid; P<.0001). Primary immunodeficiency was found in 17 (29%) of the 59 subjects tested. Thirteen other patients had secondary immune defects, and 86% of immunodeficient subjects were Hispanic or African American. CONCLUSIONS: An ICD-9-based scoring algorithm identifies patients demographically different from other hospitalized subjects who have multiple illnesses suggestive of immunodeficiency. This group contains undiagnosed minority patients with immunodeficiency.

Adolescent↗

Image processing system for interpreting motion in American Sign Language.

In this paper, an image processing algorithm is presented for the interpretation of the American Sign Language (ASL), which is one of the sign languages used by the majority of the deaf community. The process involves detection of hand motion, tracking the hand location based on the motion and classification of signs using adaptive clustering of stop positions, simple shape of the trajectory, and matching of the hand shape at the stop position.

Algorithms↗

Adaptive classification of two-dimensional gel electrophoretic spot patterns by neural networks and cluster analysis.

The interpretation of two-dimensional gel electrophoresis spot profiles can be facilitated by statistical and machine learning programs. Two different approaches to classification of spot profiles - cluster analysis and neural networks - are discussed. Neural networks for two different model patterns were designed and an algorithm for training of the net for the classification was developed. It was shown that the performance of neural networks is higher compared to cluster and principal component analysis. The possibility of combining both approaches into one process can increase reliability and speed of classification. Artificially created training sets with added random noise can be used for network training. The analysis was applied on the Streptomyces coelicolor developmental two-dimensional (2-D) gel database.

Cluster Analysis↗

A support for decision-making: cost-sensitive learning system.

This paper investigates a machine learning (ML) algorithm for supporting a decision-making system that is able to handle diagnostic problems. The input data are expressed by solved cases of patients' diagnoses, and the output is formed by a set of decision rules which may be directly exploited for a decision support. We have chosen the methodology of covering ML algorithms, namely the CN2 algorithm, as a starting point, and designed and implemented a certain extension of CN2 that comprises: advanced discretizing numerical attributes and incorporating attribute cost to economize the classification.

Algorithms↗

K-space sampling strategies.

The k-space algorithm offers a comprehensive way for classification and understanding of the imaging properties of all commonly used MR sequences. This presentation describes the basic concepts of k-space and its most relevant properties for MR imaging. The ramifications of k-space sampling is discussed for the most commonly used groups of MR sequences including gradient-echo techniques, echo-planar imaging, spin echo, and rapid acquisition relation enhanced imaging (e. g., turbo spin echo, fast spin echo). In addition, the basic problems and properties of sequences based on non-rectilinear k-space sampling, such as spiral imaging, are discussed. Their artifact behavior is significantly different from rectilinear scans, which project all imperfections along the phase-encoding directions, whereas the artifact produced by spirals are more complex and not always easily recognizable as such. An understanding of the k-space sampling offers important insight into the basic properties of a given sequence regarding signal-to-noise ratio, image distortion, resolution and contrast. It is demonstrated that the ultimate limitation in imaging speed is given by the loss of signal-to-noise ratio inherent to faster data sampling.

Algorithms↗

Management of acne.

Precise classification methods are used to define acne according to type (comedonal, papulopustular, or nodular) and severity. The relative effectiveness of several topical and systemic agents has been established in clinical trials, making possible an algorithm of specific treatment decisions based on acne classification.

Acne Vulgaris↗

Comparisons of survival predictions using survival risk ratios based on International Classification of Diseases, Ninth Revision and Abbreviated Injury Scale trauma diagnosis codes.

BACKGROUND: We conducted a comparison of methods for predicting survival using survival risk ratios (SRRs), including new comparisons based on International Classification of Diseases, Ninth Revision (ICD-9) versus Abbreviated Injury Scale (AIS) six-digit codes. METHODS: From the Pennsylvania trauma center's registry, all direct trauma admissions were collected through June 22, 1999. Patients with no comorbid medical diagnoses and both ICD-9 and AIS injury codes were used for comparisons based on a single set of data. SRRs for ICD-9 and then for AIS diagnostic codes were each calculated two ways: from the survival rate of patients with each diagnosis and when each diagnosis was an isolated diagnosis. Probabilities of survival for the cohort were calculated using each set of SRRs by the multiplicative ICISS method and, where appropriate, the minimum SRR method. These prediction sets were then internally validated against actual survival by the Hosmer-Lemeshow goodness-of-fit statistic. RESULTS: The 41,364 patients had 1,224 different ICD-9 injury diagnoses in 32,261 combinations and 1,263 corresponding AIS injury diagnoses in 31,755 combinations, ranging from 1 to 27 injuries per patient. All conventional ICD-9-based combinations of SRRs and methods had better Hosmer-Lemeshow goodness-of-fit statistic fits than their AIS-based counterparts. The minimum SRR method produced better calibration than the multiplicative methods, presumably because it did not magnify inaccuracies in the SRRs that might occur with multiplication. CONCLUSION: Predictions of survival based on anatomic injury alone can be performed using ICD-9 codes, with no advantage from extra coding of AIS diagnoses. Predictions based on the single worst SRR were closer to actual outcomes than those based on multiplying SRRs.

Abbreviated Injury Scale↗

Potential effectiveness of quality assurance screening using large but imperfect databases.

To examine the effect of imprecise classification of patient risk (severity of illness) on an otherwise highly accurate quality assurance screening technique, data on clinical outcomes were generated for a simulated hospital system consisting of 108 facilities treating approximately 565,000 patients a year. In these simulations, marked differences in facility size, casemix distribution, and quality of care were combined with random variations in outcome. Pooled data for all 108 facilities were used to create algorithms that combined 468 discrete patient risk classifications into either ten or three groups with broad, overlapping ranges of patient-specific risks of unfavorable clinical results. When derived algorithms were applied to independently generated facility-specific data, the ability to identify hospital systems with and without quality of care problems was maintained with ten, but not with three, risk groups. However, even three moderately heterogeneous risk groups were sufficient to preserve a high degree of sensitivity and specificity in screening for potential quality of care problems within individual facilities. Thus, outcome-based quality assurance screening can be highly accurate in actual health care situations in which only imprecise estimations of patient-specific risk can be achieved.

Algorithms↗

Hotelling's T2 multivariate profiling for detecting differential expression in microarrays.

The most widely used statistical methods for finding differentially expressed genes (DEGs) are essentially univariate. In this study, we present a new T(2) statistic for analyzing microarray data. We implemented our method using a multiple forward search (MFS) algorithm that is designed for selecting a subset of feature vectors in high-dimensional microarray datasets. The proposed T2 statistic is a corollary to that originally developed for multivariate analyses and possesses two prominent statistical properties. First, our method takes into account multidimensional structure of microarray data. The utilization of the information hidden in gene interactions allows for finding genes whose differential expressions are not marginally detectable in univariate testing methods. Second, the statistic has a close relationship to discriminant analyses for classification of gene expression patterns. Our search algorithm sequentially maximizes gene expression difference/distance between two groups of genes. Including such a set of DEGs into initial feature variables may increase the power of classification rules. We validated our method by using a spike-in HGU95 dataset from Affymetrix. The utility of the new method was demonstrated by application to the analyses of gene expression patterns in human liver cancers and breast cancers. Extensive bioinformatics analyses and cross-validation of DEGs identified in the application datasets showed the significant advantages of our new algorithm.

Algorithms↗

Verbal autopsies for adult deaths: their development and validation in a multicentre study.

BACKGROUND: Verbal autopsy (VA) has been widely used to ascertain causes of child deaths, but little is known about the usefulness of VA for adult deaths. This paper describes the process used to develop a VA tool for adult deaths and the results of a multicentre validation of this tool. METHODS: A mortality classification was developed by including causes of death that might be arrived at by VAs and causes that are responsive to public health interventions. An algorithm was designed for each cause in the classification, based on classifying symptoms into essential, supportive and differential. A structured questionnaire designed to elicit information on these symptoms was developed in English translated into the local languages. The tool was validated on deaths occurring at hospitals in Tanzania (315 deaths), Ethiopia (249) and Ghana (232). Hospital records of all adult deaths occurring at the study hospitals from June 1993 to April 1995 were collected prospectively. Non-medical interviewers with at least 12 years of formal education conducted VA interviews. Causes of death were diagnosed by a panel of physicians and by a computerized algorithm. The validity of the VA was assessed by comparing the VA diagnoses with hospital diagnoses. RESULTS: Specificity of VAs by physicians fell below 95% only for acute febrile illness (AFI) and TB/AIDS. Sensitivity and positive predictive value (PPV), however, varied widely both across the sites and between causes. Sensitivity was > 75% for tetanus, rabies, direct maternal causes, injuries and TB/AIDS and ranged between 60% and 74% for diarrhoea, acute abdominal conditions and AFI. The PPV was > 75% for tetanus, rabies, hepatitis and injuries and ranged between 60 and 74% for meningitis, AFI, TB/AIDS and direct maternal causes. When the communicable diseases were combined in a single group, the sensitivity was 82%, specificity 78% and PPV 85%. For the group of noncommunicable diseases the corresponding sensitivity, specificity and PPV were 71%, 87% and 67%, respectively. Use of an algorithm resulted in lower sensitivity, specificity and PPV than the VAs by physician. CONCLUSION: VAs by a panel of physicians performed better than an opinion-based algorithm. The validity of VA diagnosis was highest for AFI, direct maternal causes, TB/AIDS, tetanus, rabies and injuries.

Adult↗

Classification of single trial motor imagery EEG recordings with subject adapted non-dyadic arbitrary time-frequency tilings.

We describe a new technique for the classification of motor imagery electroencephalogram (EEG) recordings in a brain computer interface (BCI) task. The technique is based on an adaptive time-frequency analysis of EEG signals computed using local discriminant bases (LDB) derived from local cosine packets (LCP). In an offline step, the EEG data obtained from the C(3)/C(4) electrode locations of the standard 10/20 system is adaptively segmented in time, over a non-dyadic grid by maximizing the probabilistic distances between expansion coefficients corresponding to left and right hand movement imagery. This is followed by a frequency domain clustering procedure in each adapted time segment to maximize the discrimination power of the resulting time-frequency features. Then, the most discriminant features from the resulting arbitrarily segmented time-frequency plane are sorted. A principal component analysis (PCA) step is applied to reduce the dimensionality of the feature space. This reduced feature set is finally fed to a linear discriminant for classification. The online step simply computes the reduced dimensionality features determined by the offline step and feeds them to the linear discriminant. We provide experimental data to show that the method can adapt to physio-anatomical differences, subject-specific and hemisphere-specific motor imagery patterns. The algorithm was applied to all nine subjects of the BCI Competition 2002. The classification performance of the proposed algorithm varied between 70% and 92.6% across subjects using just two electrodes. The average classification accuracy was 80.6%. For comparison, we also implemented an adaptive autoregressive model based classification procedure that achieved an average error rate of 76.3% on the same subjects, and higher error rates than the proposed approach on each individual subject.

Algorithms↗

Data mining techniques for cancer detection using serum proteomic profiling.

OBJECTIVE: Pathological changes in an organ or tissue may be reflected in proteomic patterns in serum. It is possible that unique serum proteomic patterns could be used to discriminate cancer samples from non-cancer ones. Due to the complexity of proteomic profiling, a higher order analysis such as data mining is needed to uncover the differences in complex proteomic patterns. The objectives of this paper are (1) to briefly review the application of data mining techniques in proteomics for cancer detection/diagnosis; (2) to explore a novel analytic method with different feature selection methods; (3) to compare the results obtained on different datasets and that reported by Petricoin et al. in terms of detection performance and selected proteomic patterns. METHODS AND MATERIAL: Three serum SELDI MS data sets were used in this research to identify serum proteomic patterns that distinguish the serum of ovarian cancer cases from non-cancer controls. A support vector machine-based method is applied in this study, in which statistical testing and genetic algorithm-based methods are used for feature selection respectively. Leave-one-out cross validation with receiver operating characteristic (ROC) curve is used for evaluation and comparison of cancer detection performance. RESULTS AND CONCLUSIONS: The results showed that (1) data mining techniques can be successfully applied to ovarian cancer detection with a reasonably high performance; (2) the classification using features selected by the genetic algorithm consistently outperformed those selected by statistical testing in terms of accuracy and robustness; (3) the discriminatory features (proteomic patterns) can be very different from one selection method to another. In other words, the pattern selection and its classification efficiency are highly classifier dependent. Therefore, when using data mining techniques, the discrimination of cancer from normal does not depend solely upon the identity and origination of cancer-related proteins.

Biomarkers, Tumor↗

Automated particle classification based on digital acquisition and analysis of flow cytometric pulse waveforms.

In flow cytometry, the typical use of front-end analog processing limits the pulse waveform features that can be measured to pulse integral, height, and width. Direct digitizing of the waveforms provides a means for the extraction of additional features, for example, pulse skewness and kurtosis, and Fourier properties. In this work, we have first demonstrated that the Fourier properties of the pulse can be employed usefully for discrimination between different types of cells that otherwise cannot be classified by using only time-domain features of the pulse. We then implemented and evaluated automatic procedures for cell classification based on neural networks. We established that neural networks could provide an efficient means of classification of cell types without the need for user interaction. The neural networks were also employed in an innovative manner for analysis of the digital flow cytometric data without feature extraction. The performance of the neural networks was compared with that of a more conventional means of classification, the K-means clustering algorithm. Neural networks can be realized in hardware, and this, in addition to their highly parallel architecture, makes them an important potential part of real-time analysis systems. These results are discussed in terms of the design of a real-time digital data acquisition system for flow cytometry.

Animals↗

[Rough sets theory in the analysis of structure-activity relationships of quaternary quinolinium- and isoquinolinium compounds].

Relationship between chemical structure and antimicrobial activity of 72 quaternary quinolinium and isoquinolinium compounds is analyzed using the theory of rough sets. The compounds are described by 11 attributes concerning structure and are divided into 3 classes of activity. The description builds up on information system. Using the rough sets approach a smallest set of attributes significant for a high quality of classification has been found. A decision algorithm has been driven from the information system showing up important relations between structure and activity. This may be helpful in supporting decisions concerning synthesis of new antimicrobial compounds.

4-Quinolones↗

Spectral imaging perspective on cytomics.

BACKGROUND: Cytomics involves the analysis of cellular morphology and molecular phenotypes, with reference to tissue architecture and to additional metadata. To this end, a variety of imaging and nonimaging technologies need to be integrated. Spectral imaging is proposed as a tool that can simplify and enrich the extraction of morphological and molecular information. Simple-to-use instrumentation is available that mounts on standard microscopes and can generate spectral image datasets with excellent spatial and spectral resolution; these can be exploited by sophisticated analysis tools. METHODS: This report focuses on brightfield microscopy-based approaches. Cytological and histological samples were stained using nonspecific standard stains (Giemsa; hematoxylin and eosin (H&E)) or immunohistochemical (IHC) techniques employing three chromogens plus a hematoxylin counterstain. The samples were imaged using the Nuance system, a commercially available, liquid-crystal tunable-filter-based multispectral imaging platform. The resulting data sets were analyzed using spectral unmixing algorithms and/or learn-by-example classification tools. RESULTS: Spectral unmixing of Giemsa-stained guinea-pig blood films readily classified the major blood elements. Machine-learning classifiers were also successful at the same task, as well in distinguishing normal from malignant regions in a colon-cancer example, and in delineating regions of inflammation in an H&E-stained kidney sample. In an example of a multiplexed ICH sample, brown, red, and blue chromogens were isolated into separate images without crosstalk or interference from the (also blue) hematoxylin counterstain. CONCLUSION: Cytomics requires both accurate architectural segmentation as well as multiplexed molecular imaging to associate molecular phenotypes with relevant cellular and tissue compartments. Multispectral imaging can assist in both these tasks, and conveys new utility to brightfield-based microscopy approaches.

Animals↗

Whole proteome prokaryote phylogeny without sequence alignment: a K-string composition approach.

A systematic way of inferring evolutionary relatedness of microbial organisms from the oligopeptide content, i.e., frequency of amino acid K-strings in their complete proteomes, is proposed. The new method circumvents the ambiguity of choosing the genes for phylogenetic reconstruction and avoids the necessity of aligning sequences of essentially different length and gene content. The only "parameter" in the method is the length K of the oligopeptides, which serves to tune the "resolution power" of the method. The topology of the trees converges with K increasing. Applied to a total of 109 organisms, including 16 Archaea, 87 Bacteria, and 6 Eukarya, it yields an unrooted tree that agrees with the biologists' "tree of life" based on SSU rRNA comparison in a majority of basic branchings, and especially, in all lower taxa.

Algorithms↗

On the ancestral compatibility of two phylogenetic trees with nested taxa.

Compatibility of phylogenetic trees is the most important concept underlying widely-used methods for assessing the agreement of different phylogenetic trees with overlapping taxa and combining them into common supertrees to reveal the tree of life. The notion of ancestral compatibility of phylogenetic trees with nested taxa was recently introduced. In this paper we analyze in detail the meaning of this compatibility from the points of view of the local structure of the trees, of the existence of embeddings into a common supertree, and of the joint properties of their cluster representations. Our analysis leads to a very simple polynomial-time algorithm for testing this compatibility, which we have implemented and is freely available for download from the BioPerl collection of Perl modules for computational biology.

Algorithms↗