PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Numerical taxonomy of staphylococci.

Over two hundred staphylococci from human and animal sources and representatives of established species of Staphylococcus, Micrococcus and Planococcus were compared in a numerical phenetic survey using 115 unit characters. Data were analyzed using the Jaccard coefficient and the unweighted pair group method with averages algorithm. Cluster composition was not markedly affected by test error, estimated as 3.49%. The staphylococci were assigned to eighteen clusters containing four or more strains and to three single member clusters. Most of the clusters were distinct and homogeneous though two were divided into subclusters. Some of the clusters and subclusters were equated with the established taxa S. aureus, S. capitis, S. cohnii, S. epidermidis, S. haemolyticus, S. hominis, S. hyicus, S. saprophyticus, S. sciuri subspecies lentus, S. sciuri subspecies sciuri, S. simulans, S. warneri and S. xylosus, the remaining ones may represent the nuclei of additional centres of variation. The numerical data also cast doubts upon the reliability of some of the tests recommended for the identification of coagulase-negative staphylococci.

Computers↗

[Sectorization of the central 10 degrees visual field in open-angle glaucoma].

In an attempt to determine an optimal sector pattern of the central 10 degrees visual field in glaucoma, we applied the VARCLUS procedure, a new clustering algorithm provided by SAS Institute, to 379 glaucoma visual fields of the central 10-2 program of the Humphrey visual field analyzer. The subjects were 211 normal-tension glaucoma (NTG) and 168 primary open angle glaucoma (POAG) eyes with early to moderately advanced stages of visual field defects. The 68 2-degree grid test points in the central 10 degrees visual field were divided into 10 sectors. The sector pattern was compatible with the projection of nerve fiber layers and no sectors extended over the horizontal meridian. Comparison of the mean of total deviation in the sector of 76 eyes of 76 POAG patients with the maximum intraocular pressure: (IOP) > or = 25 mmHg (high-tension group) and 85 eyes of 85 NTG patients with the maximum IOP < or = 18 mmHg (low-tension group) revealed that four sectors nasal superior to the fixation were significantly more damaged in the low-tension group. We suggest that the sector pattern obtained here is useful in studying the visual field of glaucoma and that the damaging processes in the optic nervehead are not the same in the high-tension and low-tension groups.

Aged↗

A detection algorithm for multiform premature ventricular contractions.

This paper reports an algorithm developed to identify and quantify multiform PVCs. The algorithm clusters PVCs of similar morphology using a combination of time-domain and frequency-domain analysis. Initially, PVCs are grouped together on the basis of four time-domain-based morphological feature measurements. However, these time-domain-based clusters many times are nonunique because commonly encountered signal changes can cause substantial variations in the feature measurements of clinically similar beats. These redundant clusters are consolidated using two frequency-domain parameters: The First Spectral Moment (FSM) (center of gravity) of the amplitude spectrum, and the 5-Hz phase angle.

Cardiac Complexes, Premature↗

Predictions of secondary structure and solvent accessibility of the light chain of the clostridial neurotoxins.

Predictions were made of the secondary, two-dimensional (2-D) structures and side-chain solvent accessibilities of the light (L) chains of the clostridial neurotoxins (botulinum neurotoxin serotypes A-G and tetanus neurotoxin). An artificial neural network was used to make these predictions from a multiple alignment of their primary structures and was the approach used in making successful predictions for the C-fragments of these neurotoxins (Lebeda et al., J. Prot. Chem., 17:311, 1998). We also exploited the fact that the L-chains are Zn-dependent proteases. Although no other metalloproteases were found to be sequentially homologous to these neurotoxin L-chains, a sequence clustering algorithm showed that several bacterially derived Zn-dependent proteases, including thermolysin, were the most similar. A 2-D structure topology map for the type A L-chain was constructed by using thermolysin as a design template. As in thermolysin, the region containing the Zn-binding sequence motif, which is part of the active site in these neurotoxins, was predicted to be minimally solvent accessible. On the other hand, the locations of residues with highly exposed side chains were predicted to occur in non-periodic structure elements. Together, these 2-D structure and solvent accessibility predictions can be used to identify important solvent-exposed regions of the L-chain. These regions may include sites that interact with residues of the neurotoxin heavy chain, sites that bind to vesicle-docking substrates or sites that form antibody epitopes.

Algorithms↗

Mass distributed clustering: a new algorithm for repeated measurements in gene expression data.

The availability of whole-genome sequence data and high-throughput techniques such as DNA microarray enable researchers to monitor the alteration of gene expression by a certain organ or tissue in a comprehensive manner. The quantity of gene expression data can be greater than 30,000 genes per one measurement, making data clustering methods for analysis essential. Biologists usually design experimental protocols so that statistical significance can be evaluated; often, they conduct experiments in triplicate to generate a mean and standard deviation. Existing clustering methods usually use these mean or median values, rather than the original data, and take significance into account by omitting data showing large standard deviations, which eliminates potentially useful information. We propose a clustering method that uses each of the triplicate data sets as a probability distribution function instead of pooling data points into a median or mean. This method permits truly unsupervised clustering of the data from DNA microarrays.

Algorithms↗

GeneRAGE: a robust algorithm for sequence clustering and domain detection.

MOTIVATION: Efficient, accurate and automatic clustering of large protein sequence datasets, such as complete proteomes, into families, according to sequence similarity. Detection and correction of false positive and negative relationships with subsequent detection and resolution of multi-domain proteins. RESULTS: A new algorithm for the automatic clustering of protein sequence datasets has been developed. This algorithm represents all similarity relationships within the dataset in a binary matrix. Removal of false positives is achieved through subsequent symmetrification of the matrix using a Smith-Waterman dynamic programming alignment algorithm. Detection of multi-domain protein families and further false positive relationships within the symmetrical matrix is achieved through iterative processing of matrix elements with successive rounds of Smith-Waterman dynamic programming alignments. Recursive single-linkage clustering of the corrected matrix allows efficient and accurate family representation for each protein in the dataset. Initial clusters containing multi-domain families, are split into their constituent clusters using the information obtained by the multi-domain detection step. This algorithm can hence quickly and accurately cluster large protein datasets into families. Problems due to the presence of multi-domain proteins are minimized, allowing more precise clustering information to be obtained automatically. AVAILABILITY: GeneRAGE (version 1.0) executable binaries for most platforms may be obtained from the authors on request. The system is available to academic users free of charge under license.

Algorithms↗

A variable fluence step clustering and segmentation algorithm for step and shoot IMRT.

A step and shoot sequencer was developed that can be integrated into an IMRT optimization algorithm. The method uses non-uniform fluence steps and is adopted to the constraints of an MLC. It consists of a clustering, a smoothing and a segmentation routine. The performance of the algorithm is demonstrated for eight mathematical profiles of differing complexity and two optimized profiles of a clinical prostate case. The results in terms of stability, flexibility, speed and conformity fulfil the criteria for the integration into the optimization concept. The performance of the clustering routine is compared with another previously published one (Bortfeld et al 1994 Int. J. Radiat. Oncol. Biol. Ph.vs. 28 723-30) and yields slightly better results in terms of mean and maximum deviation between the optimized and the clustered protile. We discuss the specific attributes of the algorithm concerning its integration into the optimization concept.

Algorithms↗

An entropy-based algorithm for detecting clusters of cases and controls and its comparison with a method using nearest neighbours.

A new method for detecting disease clustering based on entropy is presented. For this method cases and controls are plotted on a map. The map is divided into regions. The entropy of the space is calculated as the log of the number of possible ways of placing the cases and controls in the various regions given the total number of cases and controls and the number of cases and controls in each region. The power of the entropy technique is tested against the power of the nearest neighbour technique (NNT). The entropy method is shown to be substantially more powerful than the NNT when there is more than one cluster in the space or when the clusters are near the boundary of the space.

Algorithms↗

Prediction of protein secondary structure with a reliability score estimated by local sequence clustering.

Most algorithms for protein secondary structure prediction are based on machine learning techniques, e.g. neural networks. Good architectures and learning methods have improved the performance continuously. The introduction of profile methods, e.g. PSI-BLAST, has been a major breakthrough in increasing the prediction accuracy to close to 80%. In this paper, a brute-force algorithm is proposed and the reliability of each prediction is estimated by a z-score based on local sequence clustering. This algorithm is intended to perform well for those secondary structures in a protein whose formation is mainly dominated by the neighboring sequences and short-range interactions. A reliability z-score has been defined to estimate the goodness of a putative cluster found for a query sequence in a database. The database for prediction was constructed by experimentally determined, non-redundant protein structures with <25% sequence homology, a list maintained by PDBSELECT. Our test results have shown that this new algorithm, belonging to what is known as nearest neighbor methods, performed very well within the expectation of previous methods and that the reliability z-score as defined was correlated with the reliability of prediction. This led to the possibility of making very accurate predictions for a few selected residues in a protein with an accuracy measure of Q3 > 80%. The further development of this algorithm, and a nucleation mechanism for protein folding are suggested.

Algorithms↗

An improved cluster labeling method for support vector clustering.

The support vector clustering (SVC) algorithm is a recently emerged unsupervised learning method inspired by support vector machines. One key step involved in the SVC algorithm is the cluster assignment of each data point. A new cluster labeling method for SVC is developed based on some invariant topological properties of a trained kernel radius function. Benchmark results show that the proposed method outperforms previously reported labeling techniques.

Algorithms↗

Efficient filtering methods for clustering cDNAs with spliced sequence alignment.

MOTIVATION: Clustering sequences of a full-length cDNA library into alternative splice form candidates is a very important problem. RESULTS: We developed a new efficient algorithm to cluster sequences of a full-length cDNA library into alternative splice form candidates. Current clustering algorithms for cDNAs tend to produce too many clusters containing incorrect splice form candidates. Our algorithm is based on a spliced sequence alignment algorithm that considers splice sites. The spliced sequence alignment algorithm is a variant of an ordinary dynamic programming algorithm, which requires O(nm) time for checking a pair of sequences where n and m are the lengths of the two sequences. Since the time bound is too large to perform all-pair comparison for a large set of sequences, we developed new techniques to reduce the computation time without affecting the accuracy of the output clusters. Our algorithm was applied to 21 076 mouse cDNA sequences of the FANTOM 1.10 database to examine its performance and accuracy. In these experiments, we achieved about 2-12-fold speedup against a method using only a traditional hash-based technique. Moreover, without using any information of the mouse genome sequence data or any gene data in public databases, we succeeded in listing 87-89% of all the clusters that biologists have annotated manually. AVAILABILITY: We provide a web service for cDNA clustering located at https://access.obigrid.org/ibm/cluspa/, for which registration for the OBIGrid (http://www.obigrid.org) is required.

Algorithms↗

A pairwise alignment algorithm which favors clusters of blocks.

Pairwise sequence alignments aim to decide whether two sequences are related and, if so, to exhibit their related domains. Recent works have pointed out that a significant number of true homologous sequences are missed when using classical comparison algorithms. This is the case when two homologous sequences share several little blocks of homology, too small to lead to a significant score. On the other hand, classical alignment algorithms, when detecting homologies, may fail to recognize all the significant biological signals. The aim of the paper is to give a solution to these two problems. We propose a new scoring method which tends to increase the score of an alignment when "blocks" are detected. This so-called Block-Scoring algorithm, which makes use of dynamic programming, is worth being used as a complementary tool to classical exact alignments methods. We validate our approach by applying it on a large set of biological data. Finally, we give a limit theorem for the score statistics of the algorithm.

Algorithms↗

Simultaneous feature selection and clustering using mixture models.

Clustering is a common unsupervised learning technique used to discover group structure in a set of data. While there exist many algorithms for clustering, the important issue of feature selection, that is, what attributes of the data should be used by the clustering algorithms, is rarely touched upon. Feature selection for clustering is difficult because, unlike in supervised learning, there are no class labels for the data and, thus, no obvious criteria to guide the search. Another important problem in clustering is the determination of the number of clusters, which clearly impacts and is influenced by the feature selection issue. In this paper, we propose the concept of feature saliency and introduce an expectation-maximization (EM) algorithm to estimate it, in the context of mixture-based clustering. Due to the introduction of a minimum message length model selection criterion, the saliency of irrelevant features is driven toward zero, which corresponds to performing feature selection. The criterion and algorithm are then extended to simultaneously estimate the feature saliencies and the number of clusters.

Algorithms↗

A median filter algorithm importing ISODATA dynamic clustering for medical imaging.

To improve conventional median filter algorithm employed in medical imaging, we proposed a new median filter algorithm importing ISODATA dynamic clustering for pattern recognition. The result of clustering was set as the parameters to decide if median filter was necessary and when it was, how the process was to be carried out. As shown in our test, this algorithm was capable of eliminating serious impulse noises and retain thorough image details, therefore enhanced signal to noise ratio and quality of the images in contrast with the conventional median filter algorithm.

Algorithms↗

A biologically inspired algorithm for microcalcification cluster detection.

The early detection of breast cancer greatly improves prognosis. One of the earliest signs of cancer is the formation of clusters of microcalcifications. We introduce a novel method for microcalcification detection based on a biologically inspired adaptive model of contrast detection. This model is used in conjunction with image filtering based on anisotropic diffusion and curvilinear structure removal using local energy and phase congruency. An important practical issue in automatic detection methods is the selection of parameters: we show that the parameter values for our algorithm can be estimated automatically from the image. This way, the method is made robust and essentially free of parameter tuning. We report results on mammograms from two databases and show that the detection performance can be improved by first including a normalisation scheme.

Algorithms↗

Neuropeptide and calcium-binding protein gene expression profiles predict neuronal anatomical type in the juvenile rat.

Neocortical neurones can be classified according to several independent criteria: morphological, physiological, and molecular expression (neuropeptides (NPs) and/or calcium-binding proteins (CaBPs)). While it has been suggested that particular NPs and CaBPs characterize certain anatomical subtypes of neurones, there is also considerable overlap in their expression, and little is known about simultaneous expression of multiple NPs and CaBPs in morphologically characterized neocortical neurones. Here we determined the gene expression profiles of calbindin (CB), parvalbumin (PV), calretinin (CR), neuropeptide Y (NPY), vasoactive intestinal peptide (VIP), somatostatin (SOM) and cholecystokinin (CCK) in 268 morphologically identified neurones located in layers 2-6 in the juvenile rat somatosensory neocortex. We used patch-clamp electrodes to label neurones with biocytin and harvest the cytoplasm to perform single-cell RT-multiplex PCR. Quality threshold clustering, an unsupervised algorithm that clustered neurones according to their entire profile of expressed genes, revealed seven distinct clusters. Surprisingly, each cluster preferentially contained one anatomical class. Artificial neural networks using softmax regression predicted anatomical types at nearly optimal statistical levels. Classification tree-splitting (CART), a simple binary neuropeptide decision tree algorithm, revealed the manner in which expression of the multiple mRNAs relates to different anatomical classes. Pruning the CART tree revealed the key predictors of anatomical class (in order of importance: SOM, PV, VIP, and NPY). We reveal here, for the first time, a strong relationship between specific combinations of NP and CaBP gene expressions and the anatomical class of neocortical neurones.

Animals↗

A wavelet-based algorithm for detecting clustered microcalcifications in digital mammograms.

A computerized scheme to detect clustered microcalcifications in digital mammograms has been developed. Detection of individual microcalcifications in regions of interest (ROIs) was also performed. The mammograms were previously classified into fatty and dense, according to their breast tissue. The most appropriate wavelet basis and reconstruction levels were selected. To select the wavelet basis, 40 profiles of microcalcifications were decomposed and reconstructed using different types of wavelet functions and different combinations of wavelet coefficients. The symlets with a basis of length 8 were chosen for fatty tissue. For dense tissue, the Daubechies' wavelets with a four-element basis were employed. Two methods to detect individual microcalcifications were evaluated: (a) two-dimensional wavelet transform, and (b) one-dimensional wavelet transform. The second technique yielded the best results, and was used to detect clustered microcalcifications in the complete mammogram. When detecting individual microcalcifications by using two-dimensional wavelet transform we have obtained, for fatty ROIs, a sensitivity of 71.11% at a false positive rate of 7.13 per image. For dense ROIs the sensitivity was 60.76% and the false positive rate, 7.33. The areas (A1) under the AFROC curves were 0.33+/-0.04 and 0.28+/-0.02, respectively. The one-dimensional wavelet transform method yielded 80.44% of sensitivity and 6.43 false positives per image (A1=0.39+/-0.03) for fatty ROIs, and 62.17% and 5.82 false positives per image (A1=0.37+/-0.02) for dense ROIs. For the detection of clusters of microcalcifications in the entire mammogram, the sensitivity was 80.00% with 0.94 false positives per image (A1=0.77+/-0.09) for fatty mammograms, and 72.85% of sensitivity at a false positive detection rate of 2.21 per image (A1=0.64+/-0.07) for dense mammograms. Globally, a sensitivity of 76.43% at a false positive detection rate of 1.57 per image was obtained.

Algorithms↗

Algorithm for data clustering in pattern recognition problems based on quantum mechanics.

We propose a novel clustering method that is based on physical intuition derived from quantum mechanics. Starting with given data points, we construct a scale-space probability function. Viewing the latter as the lowest eigenstate of a Schrödinger equation, we use simple analytic operations to derive a potential function whose minima determine cluster centers. The method has one parameter, determining the scale over which cluster structures are searched. We demonstrate it on data analyzed in two dimensions (chosen from the eigenvectors of the correlation matrix). The method is applicable in higher dimensions by limiting the evaluation of the Schrödinger potential to the locations of data points.

Journal Article↗