PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Rapid characterisation and identification of mycobacteria using fluorogenic enzyme tests.

Sixty representatives of selected Mycobacterium and Nocardia species were examined for their ability to cleave 79 fluorogenic synthetic enzyme substrates based on the fluorophores 7-amino-4-methylcoumarin and 4-methylumbelliferone. The resultant data were analysed using the simple matching coefficient and clustering achieved using the unweighted pair group method with arithmetic averages algorithm. Clusters corresponding to the validly described species Mycobacterium bovis, M. chelonae, M. chitae, M. farcinogenes, M. fortuitum, M. peregrinum, M. senegalense, M. smegmatis, Nocardia asteroides, and N. farcinica were circumscribed at or above the 83% similarity level. Fluorogenic probes prepared from 7-amino-4-methylcoumarin and 4-methylumbelliferone provide a rapid means of detecting taxonomically useful enzymes in small amounts of whole mycobacteria and nocardiae.

Clinical Enzyme Tests↗

OrthoMCL: identification of ortholog groups for eukaryotic genomes.

The identification of orthologous groups is useful for genome annotation, studies on gene/protein evolution, comparative genomics, and the identification of taxonomically restricted sequences. Methods successfully exploited for prokaryotic genome analysis have proved difficult to apply to eukaryotes, however, as larger genomes may contain multiple paralogous genes, and sequence information is often incomplete. OrthoMCL provides a scalable method for constructing orthologous groups across multiple eukaryotic taxa, using a Markov Cluster algorithm to group (putative) orthologs and paralogs. This method performs similarly to the INPARANOID algorithm when applied to two genomes, but can be extended to cluster orthologs from multiple species. OrthoMCL clusters are coherent with groups identified by EGO, but improved recognition of "recent" paralogs permits overlapping EGO groups representing the same gene to be merged. Comparison with previously assigned EC annotations suggests a high degree of reliability, implying utility for automated eukaryotic genome annotation. OrthoMCL has been applied to the proteome data set from seven publicly available genomes (human, fly, worm, yeast, Arabidopsis, the malaria parasite Plasmodium falciparum, and Escherichia coli). A Web interface allows queries based on individual genes or user-defined phylogenetic patterns (http://www.cbil.upenn.edu/gene-family). Analysis of clusters incorporating P. falciparum genes identifies numerous enzymes that were incompletely annotated in first-pass annotation of the parasite genome.

Animals↗

Genomic and proteomic analysis of thirty-nine structural proteins of shrimp white spot syndrome virus.

White spot syndrome virus (WSSV) virions were purified from the hemolymph of experimentally infected crayfish Procambarus clarkii, and their proteins were separated by 8 to 18% gradient sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) to give a protein profile. The visible bands were then excised from the gel, and following trypsin digestion of the reduced and alkylated WSSV proteins in the bands, the peptide sequence of each fragment was determined by liquid chromatography-nano-electrospray ionization tandem mass spectrometry (LC-nanoESI-MS/MS) using a quadrupole/time-of-flight mass spectrometer. Comparison of the resulting peptide sequence data against the nonredundant database at the National Center for Biotechnology Information identified 33 WSSV structural genes, 20 of which are reported here for the first time. Since there were six other known WSSV structural proteins that could not be identified from the SDS-PAGE bands, there must therefore be a total of at least 39 (33 + 6) WSSV structural protein genes. Only 61.5% of the WSSV structural genes have a polyadenylation signal, and preliminary analysis by 3' rapid amplification of cDNA ends suggested that some structural protein genes produced mRNA without a poly(A) tail. Microarray analysis showed that gene expression started at 2, 6, 8, 12, 18, 24, and 36 hpi for 7, 1, 4, 12, 9, 5, and 1 of the genes, respectively. Based on similarities in their time course expression patterns, a clustering algorithm was used to group the WSSV structural genes into four clusters. Genes that putatively had common or similar roles in the viral infection cycle tended to appear in the same cluster.

Amino Acid Sequence↗

Clines, clusters, and the effect of study design on the inference of human population structure.

Previously, we observed that without using prior information about individual sampling locations, a clustering algorithm applied to multilocus genotypes from worldwide human populations produced genetic clusters largely coincident with major geographic regions. It has been argued, however, that the degree of clustering is diminished by use of samples with greater uniformity in geographic distribution, and that the clusters we identified were a consequence of uneven sampling along genetic clines. Expanding our earlier dataset from 377 to 993 markers, we systematically examine the influence of several study design variables--sample size, number of loci, number of clusters, assumptions about correlations in allele frequencies across populations, and the geographic dispersion of the sample--on the "clusteredness" of individuals. With all other variables held constant, geographic dispersion is seen to have comparatively little effect on the degree of clustering. Examination of the relationship between genetic and geographic distance supports a view in which the clusters arise not as an artifact of the sampling scheme, but from small discontinuous jumps in genetic distance for most population pairs on opposite sides of geographic barriers, in comparison with genetic distance for pairs on the same side. Thus, analysis of the 993-locus dataset corroborates our earlier results: if enough markers are used with a sufficiently large worldwide sample, individuals can be partitioned into genetic clusters that match major geographic subdivisions of the globe, with some individuals from intermediate geographic locations having mixed membership in the clusters that correspond to neighboring regions.

Alleles↗

A New Method for Database Searching and Clustering.

An iterative database searching method is introduced and applied to the design of a database clustering procedure. The search method virtually never produces false positive hits while determining meaningfully large sets of sequences related to the query. A novel set-theoretic database clustering algorithm exploits this feature and avoids a traditional, distance-based clustering step. This makes it fast and applicable to data-sets of the size of, e.g., the Swiss-Prot database. In practice we achieve unambiguous assignment of 80% of Swiss-Prot sequences to non-overlapping sequence clusters in an entirely automatic fashion.

Journal Article↗

Profile clusters in the WAIS-R standardization sample.

In this study we applied clustering procedures to a subgroup of 341 participants from the WAIS-R standardization sample. These individuals were selected by virtue of their having full-scale profiles characterized by scatter of greater than 8 scaled scores. Using a hierarchical clustering algorithm, a multistage procedure was used to establish and evaluate a cluster solution. The subject selection and clustering methods were successful in revealing a set of 9 profile types characterized by unique profile shapes. All profiles were associated with FSIQs that were at least in the average range. Seven of the profiles were characterized by specific subtest strengths, only 1 with subtest weaknesses. Examination of the external correlates of profile membership revealed differences among profile groups for age, marital status, education, and occupation. Our findings suggest that variability in and across the 9 profile types in this sample reflects increased contributions of unique abilities in comparison to the influence of the underlying primary and secondary WAIS-R dimensions of cognitive ability.

Adolescent↗

Local linear independent component analysis based on clustering.

In standard Independent Component Analysis (ICA), a linear data model is used for a global description of the data. Even though linear ICA yields meaningful results in many cases, it can provide a crude approximation only for general nonlinear data distributions. In this paper a new structure is proposed, where local ICA models are used in connection with a suitable grouping algorithm clustering the data. The clustering part is responsible for an overall coarse nonlinear representation of the data, while linear ICA models of each cluster are used for describing local features of the data. The goal is to represent the data better than in linear ICA while avoiding computational difficulties related with nonlinear ICA. Several data grouping methods are considered, including standard K-means clustering, self-organizing maps, and neural gas. Connections to existing methods are discussed, and experimental results are given for artificial data and natural images. Furthermore, a general theoretical framework encompassing a large number of methods for representing data is introduced. These range from global, dense representation methods to local, very sparse coding methods. The proposed local ICA methods lie between these two extremes.

Algorithms↗

Efficiently mining gene expression data via a novel parameterless clustering method.

Clustering analysis has been an important research topic in the machine learning field due to the wide applications. In recent years, it has even become a valuable and useful tool for in-silico analysis of microarray or gene expression data. Although a number of clustering methods have been proposed, they are confronted with difficulties in meeting the requirements of automation, high quality, and high efficiency at the same time. In this paper, we propose a novel, parameterless and efficient clustering algorithm, namely, Correlation Search Technique (CST), which fits for analysis of gene expression data. The unique feature of CST is it incorporates the validation techniques into the clustering process so that high quality clustering results can be produced on the fly. Through experimental evaluation, CST is shown to outperform other clustering methods greatly in terms of clustering quality, efficiency, and automation on both of synthetic and real data sets.

Algorithms↗

A data mining approach to the development of a diagnostic test for male infertility.

The paper presents a database of published Y chromosome deletions and the results of analyzing the database with data mining and other heuristic techniques with the goal of developing a diagnostic test for male infertility. The database describes 382 patients for which 177 markers were tested. Two data mining techniques, clustering and decision tree induction were used, as well as a heuristic set cover algorithm. Clustering was used to group markers according to their appearance across patients, while a heuristic set covering algorithm was used to select as small a set of markers that cover as many patients with deletions as possible. This algorithm created a diagnostic set of 13 markers that cover more than 90% of the patients with deletions. Finally, decision tree induction was used to relate deletion patterns to the severity of the clinical phenotype. A decision tree induced from the data uses 5 markers, all of which are also in the diagnostic set of 13 markers, to show relations between the severity of the clinical phenotype and deletion patterns which have not been known previously.

Algorithms↗

Bayesian clustering using hidden Markov random fields in spatial population genetics.

We introduce a new Bayesian clustering algorithm for studying population structure using individually geo-referenced multilocus data sets. The algorithm is based on the concept of hidden Markov random field, which models the spatial dependencies at the cluster membership level. We argue that (i) a Markov chain Monte Carlo procedure can implement the algorithm efficiently, (ii) it can detect significant geographical discontinuities in allele frequencies and regulate the number of clusters, (iii) it can check whether the clusters obtained without the use of spatial priors are robust to the hypothesis of discontinuous geographical variation in allele frequencies, and (iv) it can reduce the number of loci required to obtain accurate assignments. We illustrate and discuss the implementation issues with the Scandinavian brown bear and the human CEPH diversity panel data set.

Animals↗

CLUE: cluster-based retrieval of images by unsupervised learning.

In a typical content-based image retrieval (CBIR) system, target images (images in the database) are sorted by feature similarities with respect to the query. Similarities among target images are usually ignored. This paper introduces a new technique, cluster-based retrieval of images by unsupervised learning (CLUE), for improving user interaction with image retrieval systems by fully exploiting the similarity information. CLUE retrieves image clusters by applying a graph-theoretic clustering algorithm to a collection of images in the vicinity of the query. Clustering in CLUE is dynamic. In particular, clusters formed depend on which images are retrieved in response to the query. CLUE can be combined with any real-valued symmetric similarity measure (metric or nonmetric). Thus, it may be embedded in many current CBIR systems, including relevance feedback systems. The performance of an experimental image retrieval system using CLUE is evaluated on a database of around 60,000 images from COREL. Empirical results demonstrate improved performance compared with a CBIR system using the same image similarity measure. In addition, results on images returned by Google's Image Search reveal the potential of applying CLUE to real-world image data and integrating CLUE as a part of the interface for keyword-based image retrieval systems.

Algorithms↗

Identification of simple sequence repeats (SSRs) in olive ( Olea europaea L.).

A small insert genomic library of Olea europaea L., highly enriched in (GA/CT) n repeats, was obtained using the procedure of Kandpal et al. (1994). The sequencing of 103 clones randomly extracted from this library allowed the identification of 56 unique genomic inserts containing simple sequence repeat regions made by at least three single repeats. A sample of 20 primer pairs out of the 42 available were tested for functionality using the six olive varieties whose DNA served for library construction. All primer pairs succeeded in amplifying at least one product from the six DNA samples, and ten pairs detecting more than one allele were used for the genetic characterisation of a panel of 20 olive accessions belonging to 16 distinct varieties. A total of 57 alleles were detected among the 20 genotypes at the ten polymorphic SSR loci. The remaining primer pair allowed the amplification of a single SSR allele for all accessions plus a longer fragment for some genotypes. Considering the simple sequence repeat polymorphism, 5.7 alleles were scored on average for each of the ten SSR loci. A genetic dissimilarity matrix, based on the proportion of shared alleles among all the pair-wise combinations of genotypes, was constructed and used to disentangle the genetic relationships among varieties by means of the UPGMA clustering algorithm. Graphical representation of the results showed the presence of two distinct clusters of varieties. The first cluster grouped the varieties cultivated on the Ionian Sea coasts. The second cluster showed two subdivisions: the first sub-cluster agglomerated the varieties from some inland areas of Calabria; the second grouped the remaining varieties from Basilicata and Apulia cultivated in nearby areas. Results of cluster analysis showed a significant relationship between the multilocus genetic similarities and the geographic origin of the cultivars.

Journal Article↗

Correlation between strand asymmetry and phylogeny in mitochondrial DNA.

An evolutionary distance is introduced in order to propose an efficient and feasible procedure for phylogeny studies. Our analysis are based on the strand asymmetry property of mitochondrial DNA, but can be applied to other genomes. Comparison of our results with those reported in conventional phylogenetic trees, gives confidence about our approximation. Our findings support the hypotheses about the origin of the skew and its dependence upon evolutionary pressures, and improves previous efforts on using the strand asymmetry property of genomes for phylogeny inference. For the evolutionary distance introduced here, we observe that the more adequate technique for tree reconstructions correspond to an average link method which employs a sequential clustering algorithm.

Algorithms↗

Numerical classification of sporoactinomycetes containing meso-diaminopimelic acid in the cell wall.

One hundred and thirty actinomycetes representing 19 genera and 50 species were compared in a numerical phenetic survey using 108 unit characters. Data were examined using the simple matching (SSM), Jaccard (SJ) and pattern (DP) coefficients and clustering was achieved using both the single and unweighted pair group average algorithms. Cluster composition was barely affected by the statistics used or by test error, estimated at 2.1%. Over 80% of the strains were assigned to 2 clusters containing between two and 25 organisms. Most of the clusters were distinct and homogeneous though two were divided into subclusters. Some of the clusters and subclusters were equated with the established taxa Actinomadura madurae, Actinomadura pelletieri, Dermatophilus congolensis, Geodermatophilus obscurus, Microbispora spp., Micromonospora spp., Micropolyspora brevicatena, Micropolyspora faeni, Nocardia spp., Nocardiopsis (Actinomadura) dassonvillei, Planobispora spp., Planomonospora spp., Saccharomonospora viridis, Streptomyces somaliensis, Thermoactinomyces candidus, Thermoactinomyces dichotomica, Thermoactinomyces sacchari, Thermoactinomyces vulgaris and "Thermomonospora fusca'. The numerical data, together with results from previous chemical and genetical studies, provide sufficient evidence for the transfer of Micropolyspora brevicatena to Nocardia as Nocardia brevicatena comb. nov.

Actinomycetales↗

Image analysis and quantification of atherosclerosis using MRI.

This paper describes an image processing, pattern recognition, and computer graphics system for the noninvasive identification and evaluation of atherosclerosis using multidimensional Magnetic Resonance Imaging (MRI). Particular emphasis has been placed on the problem of developing a pattern recognition system for noninvasively identifying the different plaque classes involved in atherosclerosis using minimal a priori information. This pattern recognition technique involves an extension of the ISODATA clustering algorithm to include an information theoretic criterion (Consistent Akaike Information Criterion) to provide a measure of the fit of the cluster composition at a particular iteration to the actual data. A rapid 3-D display system is also described for the simultaneous display of multiple data classes resulting from the tissue identification process. This work demonstrates the feasibility of developing a "high information content" display which will aid in the diagnosis and analysis of the atherosclerotic disease process. Such capability will permit detailed and quantitative studies to assess the effectiveness of therapies, such as drug, exercise, and dietary regimens.

Algorithms↗

High-speed face recognition based on discrete cosine transform and RBF neural networks.

In this paper, an efficient method for high-speed face recognition based on the discrete cosine transform (DCT), the Fisher's linear discriminant (FLD) and radial basis function (RBF) neural networks is presented. First, the dimensionality of the original face image is reduced by using the DCT and the large area illumination variations are alleviated by discarding the first few low-frequency DCT coefficients. Next, the truncated DCT coefficient vectors are clustered using the proposed clustering algorithm. This process makes the subsequent FLD more efficient. After implementing the FLD, the most discriminating and invariant facial features are maintained and the training samples are clustered well. As a consequence, further parameter estimation for the RBF neural networks is fulfilled easily which facilitates fast training in the RBF neural networks. Simulation results show that the proposed system achieves excellent performance with high training and recognition speed, high recognition rate as well as very good illumination robustness.

Algorithms↗

Automatic clustering of orthologs and inparalogs shared by multiple proteomes.

MOTIVATION: The complete sequencing of many genomes has made it possible to identify orthologous genes descending from a common ancestor. However, reconstruction of evolutionary history over long time periods faces many challenges due to gene duplications and losses. Identification of orthologous groups shared by multiple proteomes therefore becomes a clustering problem in which an optimal compromise between conflicting evidences needs to be found. RESULTS: Here we present a new proteome-scale analysis program called MultiParanoid that can automatically find orthology relationships between proteins in multiple proteomes. The software is an extension of the InParanoid program that identifies orthologs and inparalogs in pairwise proteome comparisons. MultiParanoid applies a clustering algorithm to merge multiple pairwise ortholog groups from InParanoid into multi-species ortholog groups. To avoid outparalogs in the same cluster, MultiParanoid only combines species that share the same last ancestor. To validate the clustering technique, we compared the results to a reference set obtained by manual phylogenetic analysis. We further compared the results to ortholog groups in KOGs and OrthoMCL, which revealed that MultiParanoid produces substantially fewer outparalogs than these resources. AVAILABILITY: MultiParanoid is a freely available standalone program that enables efficient orthology analysis much needed in the post-genomic era. A web-based service providing access to the original datasets, the resulting groups of orthologs, and the source code of the program can be found at http://multiparanoid.cgb.ki.se.

Algorithms↗

CAGNet: a structure-aware clustering-alternated graph network for cell-cell interaction inference in spatial transcriptomics.

MOTIVATION: Understanding cell-cell interactions (CCIs) in spatial transcriptomics is crucial for uncovering the spatial organization and functional heterogeneity of tissues. However, existing graph-based models typically rely on static clustering or fixed adjacency structures, which limits their ability to capture dynamic cellular relationships. RESULTS: We propose CAGNet, a two-stage framework for CCI inference from spatial transcriptomics data. In Stage 1, a Graph Attention Network encoder with joint feature and graph reconstruction learns structure-aware node embeddings from spatial gene expression profiles. In Stage 2, an alternating optimization mechanism iteratively updates cluster centers via KL-guided soft assignment and refines node embeddings through spatial graph reconstruction, establishing a closed-loop between representation learning and clustering. Experiments on three 10x Genomics Visium datasets demonstrate that CAGNet consistently outperforms six CCI inference baselines across ACC, AUC, AP, Precision, Recall, and F1. CAGNet also achieves the highest Adjusted Rand Index on all three datasets against six spatial domain identification methods, confirming that the learned embeddings capture biologically relevant spatial organization. Information-theoretic analysis further shows that CAGNet retains the highest mutual information between input features and learned embeddings among all compared methods. Ablation studies and 5-fold cross-validation confirm the contribution of each component and the reproducibility of the results. AVAILABILITY: The proposed method is implemented in the CAGNet package available at http://github.com/mahan1233333-maker/CAGNet .

Spatial Transcriptomics↗