PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Cluster analysis in family psychology research.

This article discusses the use of cluster analysis in family psychology research. It provides an overview of potential clustering methods, the steps involved in cluster analysis, hierarchical and nonhierarchical clustering methods, and validation and interpretation of cluster solutions. The article also reviews 5 uses of clustering in family psychology research: (a) deriving family types, (b) studying families over time, (c) as an interface between qualitative and quantitative methods, (d) as an alternative to multivariate interactions in linear models, and (e) as a data reduction technique for small samples. The article concludes with some cautions for using clustering in family psychology research.

Algorithms↗

Text mining of DNA sequence homology searches.

Primary tasks in analysis and annotation of expressed sequence tag (EST) datasets are to identify similarity among sequences by unsupervised clustering and assign putative function based on BLAST homology searches. We investigated the usefulness of text mining as a simple approach for further higher-level clustering of EST datasets using IBM Intelligent Miner for Text v2.3 tools. Agglomerative and k-means clustering tools were used to cluster BLASTx homology search documents from two onion EST datasets and optimised by pre-processing and pruning. Subjective evaluation confirmed that these tools provided biologically useful and complementary views of the two libraries, provided new insights into their composition and revealed clusters previously identified by human experts. We compared BLASTx textual clusters for two gene families with their DNA sequence-based clusters and confirmed that these shared similar morphology.

Abstracting and Indexing↗

Profiling local optima in K-means clustering: developing a diagnostic technique.

Using the cluster generation procedure proposed by D. Steinley and R. Henson (2005), the author investigated the performance of K-means clustering under the following scenarios: (a) different probabilities of cluster overlap; (b) different types of cluster overlap; (c) varying samples sizes, clusters, and dimensions; (d) different multivariate distributions of clusters; and (e) various multidimensional data structures. The results are evaluated in terms of the Hubert-Arabie adjusted Rand index, and several observations concerning the performance of K-means clustering are made. Finally, the article concludes with the proposal of a diagnostic technique indicating when the partitioning given by a K-means cluster analysis can be trusted. By combining the information from several observable characteristics of the data (number of clusters, number of variables, sample size, etc.) with the prevalence of unique local optima in several thousand implementations of the K-means algorithm, the author provides a method capable of guiding key data-analysis decisions.

Algorithms↗

Quantitative structure-antitumor activity relationships of camptothecin analogues: cluster analysis and genetic algorithm-based studies.

Topoisomerase 1 (top1) inhibitors are proving useful against a range of refractory tumors, and there is considerable interest in the development of additional top1 agents. Despite crystallographic studies, the binding site and ligand properties that lead to activity are poorly understood. Here we report a unique approach to quantitative structure-activity relationship (QSAR) analysis based on the National Cancer Institute's (NCI) drug databases. In 1990, the NCI established a drug discovery program in which compounds are tested for their ability to inhibit the growth of 60 different human cancer cell lines in culture. More than 70 000 compounds have been screened, and patterns of activity against the 60 cell lines have been found to encode rich information on mechanisms of drug action and drug resistance. Here, we use hierarchical clustering to define antitumor activity patterns in a data set of 167 tested camptothecins (CPTs) in the NCI drug database. The average pairwise Pearson correlation coefficient between activity patterns for the CPT set was 0.70. Coherence between chemical structures and their activity patterns was observed. QSAR studies were carried out using the mean 50% growth inhibitory concentrations (GI(50)) for 60 cell lines as the dependent variables. Different statistical methods, including stepwise linear regression, principal component regression (PCR), partial least-squares regression (PLS), and fully cross-validated genetic function approximation (GFA) were applied to construct quantitative structure-antitumor relationship models. For our data set, the GFA method performed better in terms of correlation coefficients and cross-validation analysis. A number of molecular descriptors were identified as being correlated with antitumor activity. Included were partial atomic charges and three interatomic distances that define the relative spatial dispositions of three significant atoms (the hydroxyl hydrogen of the E-ring, the lactone carbonyl oxygen of the E-ring, and the carbonyl oxygen of the D-ring). The cross-validated r(2) for the final GFA model was 0.783, indicating a predictive QSAR model.

Algorithms↗

Developing an empirical typology for regular exercise.

BACKGROUND: Tailored interventions require the identification of distinct homogenous subgroups that will benefit from different intervention materials. One way to identify such subgroups is to use cluster analysis to identify an empirical typology. METHODS: A sample of 346 adults completed surveys through a telephone interview that included questions related to participating in regular exercise. The three variables used in the cluster analysis were the Pros of Exercise, the Cons of Exercise, and Exercise Self-Efficacy. RESULTS: Six resulting clusters were labeled Disengaged, Immotive, Relapse Risk, Early Action, Maintainers, and Habituated. A series of analyses tested the internal and external validity of the typology. The internal validity test revealed that four of the clusters demonstrated high stability and replicability, while the Relapse Risk and Early Action clusters were less stable. Differences among clusters on self-reported exercise behavior and a strong association with stage of change for regular exercise provided external validity evidence of the typology. CONCLUSIONS: The resulting typology reflects a range of motivational patterns that are likely to be responsive to different types of messages and strategies regarding adoption and maintenance of regular exercise. The typology also generates a number of hypotheses about the identified clusters that can be empirically tested in further studies.

Adult↗

Novel technique for preprocessing high dimensional time-course data from DNA microarray: mathematical model-based clustering.

MOTIVATION: Classifying genes into clusters depending on their expression profiles is one of the most important analysis techniques for microarray data. Because temporal gene expression profiles are indicative of the dynamic functional properties of genes, the application of clustering analysis to time-course data allows the more precise division of genes into functional classes. Conventional clustering methods treat the sampling data at each time point as data obtained under different experimental conditions without considering the continuity of time-course data between time periods t and t+1. Here, we propose a method designated mathematical model-based clustering (MMBC). RESULTS: The proposed method, designated MMBC, was applied to artificial data and time-course data obtained using Saccharomyces cerevisiae. Our method is able to divide data into clusters more accurately and coherently than conventional clustering methods. Furthermore, MMBC is more tolerant to noise than conventional clustering methods. AVAILABILITY: Software is available upon request. CONTACT: taizo@brs.kyushu-u.ac.jp.

Algorithms↗

Fuzzy C-means method for clustering microarray data.

MOTIVATION: Clustering analysis of data from DNA microarray hybridization studies is essential for identifying biologically relevant groups of genes. Partitional clustering methods such as K-means or self-organizing maps assign each gene to a single cluster. However, these methods do not provide information about the influence of a given gene for the overall shape of clusters. Here we apply a fuzzy partitioning method, Fuzzy C-means (FCM), to attribute cluster membership values to genes. RESULTS: A major problem in applying the FCM method for clustering microarray data is the choice of the fuzziness parameter m. We show that the commonly used value m = 2 is not appropriate for some data sets, and that optimal values for m vary widely from one data set to another. We propose an empirical method, based on the distribution of distances between genes in a given data set, to determine an adequate value for m. By setting threshold levels for the membership values, genes which are tigthly associated to a given cluster can be selected. Using a yeast cell cycle data set as an example, we show that this selection increases the overall biological significance of the genes within the cluster. AVAILABILITY: Supplementary text and Matlab functions are available at http://www-igbmc.u-strasbg.fr/fcm/

Algorithms↗

Using cluster analysis for medical resource decision making.

Escalating costs of health care delivery have in the recent past often made the health care industry investigate, adapt, and apply those management techniques relating to budgeting, resource control, and forecasting that have long been used in the manufacturing sector. A strategy that has contributed much in this direction is the definition and classification of a hospital's output into "products" or groups of patients that impose similar resource or cost demands on the hospital. Existing classification schemes have frequently employed cluster analysis in generating these groupings. Unfortunately, the myriad articles and books on clustering and classification contain few formalized selection methodologies for choosing a technique for solving a particular problem, hence they often leave the novice investigator at a loss. This paper reviews the literature on clustering, particularly as it has been applied in the medical resource-utilization domain, addresses the critical choices facing an investigator in the medical field using cluster analysis, and offers suggestions (using the example of clustering low-vision patients) for how such choices can be made.

Algorithms↗

Functional clustering of yeast proteins from the protein-protein interaction network.

BACKGROUND: The abundant data available for protein interaction networks have not yet been fully understood. New types of analyses are needed to reveal organizational principles of these networks to investigate the details of functional and regulatory clusters of proteins. RESULTS: In the present work, individual clusters identified by an eigenmode analysis of the connectivity matrix of the protein-protein interaction network in yeast are investigated for possible functional relationships among the members of the cluster. With our functional clustering we have successfully predicted several new protein-protein interactions that indeed have been reported recently. CONCLUSION: Eigenmode analysis of the entire connectivity matrix yields both a global and a detailed view of the network. We have shown that the eigenmode clustering not only is guided by the number of proteins with which each protein interacts, but also leads to functional clustering that can be applied to predict new protein interactions.

Algorithms↗

Database clustering with a combination of fingerprint and maximum common substructure methods.

We present an efficient method to cluster large chemical databases in a stepwise manner. Databases are first clustered with an extended exclusion sphere algorithm based on Tanimoto coefficients calculated from Daylight fingerprints. Substructures are then extracted from clusters by iterative application of a maximum common substructure algorithm. Clusters with common substructures are merged through a second application of an exclusion sphere algorithm. In a separate step, singletons are compared to cluster substructures and added to a cluster if similarity is sufficiently high. The method identifies tight clusters with conserved substructures and generates singletons only if structures are truly distinct from all other library members. The method has successfully been applied to identify the most frequently occurring scaffolds in databases, for the selection of analogues of screening hits and in the prioritization of chemical libraries offered by commercial vendors.

Journal Article↗

Classification of malignant and benign tumors using boundary characteristics in breast ultrasonograms.

We evaluated various spiculate and jagged margin shape features. These are known to be malignant characteristics in breast sonograms. A total of 79 breast ultrasonograms (60 benign, 19 malignant) containing solid breast nodules were evaluated. To determine the boundary of lesions, Markov random field segmentation was used. Our goal was to classify benign and malignant lesions on the breast sonogram. Our algorithm consisted of two steps: segmentation and classification. In the first step, a breast sonogram was segmented using low resolution and Gaussian-Markov random field. The fuzzy clustering method algorithm was then applied to the preprocessed image to initialize the segmentation. Next, to discriminate benign and malignant tumors three types of lesion characteristics were investigated: jag count, compactness, and acutance. Jag count was calculated based on the derivative of curvature, acutance was defined as gray-level variations across the lesion boundary, and compactness was defined as the ratio of boundary complexity to the enclosed area. Sensitivity of the three boundary features (jag count, compactness, and acutance) was 95.1, 94.1, and 81.1%, respectively, and their specificities were 97.2, 92.0, and 78%, respectively. The jag count performed best among the three boundary features. Our results indicate that computerized analysis of boundary characteristics can be an effective method for classifying solid breast nodules in ultrasonograms as malignant or benign. We found that curvature analysis was the best shape features. The curvature method classifies better than the compactness or acutance methods.

Breast Neoplasms↗

Multi-class clustering and prediction in the analysis of microarray data.

DNA microarray technology provides tools for studying the expression profiles of a large number of distinct genes simultaneously. This technology has been applied to sample clustering and sample prediction. Because of a large number of genes measured, many of the genes in the original data set are irrelevant to the analysis. Selection of discriminatory genes is critical to the accuracy of clustering and prediction. This paper considers statistical significance testing approach to selecting discriminatory gene sets for multi-class clustering and prediction of experimental samples. A toxicogenomic data set with nine treatments (a control and eight metals, As, Cd, Ni, Cr, Sb, Pb, Cu, and AsV with a total of 55 samples) is used to illustrate a general framework of the approach. Among four selected gene sets, a gene set omega(I) formed by the intersection of the F-test and the set of the union of one-versus-all t-tests performs the best in terms of clustering as well as prediction. Hierarchical and two modified partition (k-means) methods all show that the set omega(I) is able to group the 55 samples into seven clusters reasonably well, in which the As and AsV samples are considered as one cluster (the same group) as are the Cd and Cu samples. With respect to prediction, the overall accuracy for the gene set omega(I) using the nearest neighbors algorithm to predict 55 samples into one of the nine treatments is 85%.

Algorithms↗

Analysis of selection methodologies for combinatorial library design.

We have implemented and adapted in Pralins (Program for Rational Analysis of Libraries in silico), the most popular sparse (cherry picking) and full array (sublibrary) selection algorithms: hierarchical clustering, k-means clustering, Optimum Binning, Jarvis Patrick, Pral-SE (partitioning techniques) and MaxSum, MaxMin, MaxMin averaged, DN2, CTD (distance-based methods). We have validated the program with an already synthesized three-component combinatorial library of FXR partial agonists characterized by standard computational chemistry descriptors as case study. This has let us analyze the goodness of both the partitioning techniques for space division and all the selection methodologies with respect to representativity in terms of population and space coverage for different selection sizes. Within the chemical space analyzed, both hierarchical clustering and Optimum Binning division strategies are found to be the most advantageous reference space divisions to be used in the subsequent population and space coverage studies. Complete hierarchical clustering appears also to be the preferred selection methodology for both sparse and full array problems. The full array restriction fulfillment can easily be overcome by convenient optimization algorithms that allow optimal reagent selection preserving > 90% of the population coverage.

Combinatorial Chemistry Techniques↗

scFANCL: Dual contrastive learning with false-negative correction at cell level for single-cell RNA-seq clustering.

BACKGROUND: Single-cell RNA sequencing (scRNA-seq) enables cellular characterization at single-cell resolution. However, its high dimensionality, sparsity, and noise make clustering challenging. Approaches utilizing contrastive learning and data augmentation have been introduced to improve representation quality for scRNA-seq clustering. In particular, dual contrastive frameworks combining instance- and cluster-level objectives can capture both cell-cell similarities and inter-cluster variations. However, existing dual contrastive frameworks focus primarily on discrete cluster boundaries, neglecting the biological continuity inherent in scRNA-seq data. METHODS: We propose scFANCL, a dual contrastive framework designed to capture biological continuity in scRNA data. Rather than treating all non-augmented samples as negatives, scFANCL applies a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, preserving continuous transcriptional relationships among them while maintaining inter-cluster separation. RESULTS: Extensive experiments across seven publicly available scRNA-seq datasets demonstrated that scFANCL achieves competitive clustering performance compared with existing baseline methods, consistently yielding high ARI and NMI scores across datasets of varying size and complexity. Ablation studies further confirmed the contribution of the false negative filtering component, showing measurable improvements over variants without filtering. Downstream analyses further suggest that the learned embeddings may reflect biologically meaningful transcriptional transitions, including continuous differentiation trajectories within related cell types. The source code is available at https://github.com/mjuailab/scFANCL . CONCLUSIONS: scFANCL addresses a key limitation of conventional contrastive learning by applying a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, thereby preserving biological continuity within cell types while maintaining inter-cluster separation. Evaluations across seven benchmark scRNA-seq datasets demonstrate competitive clustering performance, with learned embeddings capturing biologically meaningful transcriptional structure and characteristics of rare cell populations.

Clustering Algorithms↗

A comprehensive set of protein complexes in yeast: mining large scale protein-protein interaction screens.

MOTIVATION: The analysis of protein-protein interactions allows for detailed exploration of the cellular machinery. The biochemical purification of protein complexes followed by identification of components by mass spectrometry is currently the method, which delivers the most reliable information--albeit that the data sets are still difficult to interpret. Consolidating individual experiments into protein complexes, especially for high-throughput screens, is complicated by many contaminants, the occurrence of proteins in otherwise dissimilar purifications due to functional re-use and technical limitations in the detection. A non-redundant collection of protein complexes from experimental data would be useful for biological interpretation, but manual assembly is tedious and often inconsistent. RESULTS: Here, we introduce a measure to define similarity within collections of purifications and generate a set of minimally redundant, comprehensive complexes using unsupervised clustering. AVAILABILITY: Programs and results are freely available from http://www.bork.embl-heidelberg.de/Docu/purclust/

Algorithms↗

cluML: A markup language for clustering and cluster validity assessment of microarray data.

cluML is a new markup language for microarray data clustering and cluster validity assessment. The XML-based format has been designed to address some of the limitations observed in traditional formats, such as inability to store multiple clustering (including biclustering) and validation results within a dataset. cluML is an effective tool to support biomedical knowledge representation in gene expression data analysis. Although cluML was developed for DNA microarray analysis applications, it can be effectively used for the representation of clustering and for the validation of other biomedical and physical data that has no limitations.

Algorithms↗

Automated texture-based segmentation of ultrasound images of the prostate.

Segmenting two-dimensional images of the prostate into prostate and nonprostate regions is required when forming a three-dimensional image of the prostate from a set of parallel two-dimensional images. The texture-based segmentation method presented here is a pixel classifier based on four texture energy measures associated with each pixel in the image. An automated clustering procedure is used to label each pixel in the image with the label of its most probable class. The segmented images produced as the result of applying the algorithm to an example image are presented and discussed. The automated segmentation algorithm has been found to hold promise as an automated segmentation method.

Adult↗

Molecular dynamics simulations of beta-turn forming tetra- and hexapeptides.

It was previously shown that the structural ensemble of model peptides DDKG and GKDG (H. Ishii et al. Biopolymers 24, 2045-2056, 1985), DEKS (A. Otter et al. J. Biomol. Struct. Dyn. 7, 455-476, 1989) NPGQ (F. R. Carbone et al. Int. J. Pept. Protein. Res. 26, 498-508, 1985), SALN (H. Santa et al. J. Biomol. Struct. Dyn. 16, 1033-1041, 1999), SYPFDV and SYPYDV (J. Yao et al. J. Mol. Biol. 243, 736-753, 1994), VP(D)AH and VP(D)SH (B. Imperiali et al. J. Am. Chem. Soc. 114, 3182-3188, 1992) in solution contains a significant - or in some cases dominant - proportion of beta-turn conformation. In this study, a protein database was searched for the above, unprotected sequences which incorporate only L-amino acid residues. Simulated annealing and 25 ns MD simulations of structures were also performed. The DSSP and STRIDE secondary structure-assigning algorithms and clustering were used to analyze trajectories and i, i+3 hydrogen bonds were also sought. The DSSP analysis showed a fluctuation between beta-turn and random meander structure, although bend structures were not detected because of the insufficient length of peptide chains. This alternating trend was confirmed when the STRIDE algorithm was used to analyze trajectories, but STRIDE assigned more turn structures. The population of the strongest clusters was above 40% and the middle structures adopted beta-turn structure for most sequences. These results are in good agreement with previous experimental results and support the idea of the ultra-marginal stability of turns in the absence of stabilizing long-range interactions of the neighboring segments of a polypeptide chain. However, interactions between the side-chains in tetrapeptides could also contribute to turn stability and result in unusual stability in some cases. Our observations suggest that such interactions are the consequence rather than the driving force of turn formation.

Amino Acid Sequence↗