PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Algorithms for automated characterization of cell populations in thick specimens from 3-D confocal fluorescence microscopy data.

Methods are presented for the automated, quantitative and three-dimensional (3-D) analysis of cell populations in thick, essentially intact tissue sections while maintaining intercell spatial relationships. This analysis replaces current manual methods which are tedious and subjective. The thick sample is imaged in three dimensions using a confocal scanning laser microscope. The stack of optical slices is processed by a 3-D segmentation algorithm that separates touching and overlapping structures using localization constraints. Adaptive data reduction is used to achieve computational efficiency. A hierarchical cluster analysis algorithm is used automatically to characterize the cell population by a variety of cell features. It allows automatic detection and characterization of patterns such as the 3-D spatial clustering of cells, and the relative distributions of cells of various sizes. It also permits the detection of structures that are much smaller, larger, brighter, darker, or differently shaped than the rest of the population. The overall method is demonstrated for a set of rat brain tissue sections that were labelled for tyrosine hydroxylase using fluorescein-conjugated antibodies. The automated system was verified by comparison with computer-assisted manual counts from the same image fields.

Algorithms↗

An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer.

This study presents a novel integrative approach for the analysis of high-dimensional gene expression data, leveraging the complementary strengths of unsupervised clustering and supervised classification. Using K-means clustering, the data set is stratified into three distinct clusters, revealing intrinsic biological patterns and relationships. The resulting cluster assignments are subsequently used as pseudolabels to train machine learning models, including support vector machines, random forest, and a stacking ensemble classifier. To validate and enhance the robustness of clustering, complementary methods, such as hierarchical clustering and density-based spatial clustering of applications with noise (DBSCAN), are used, with results visualized through principal component analysis-driven dimensionality reduction. The high predictive accuracy achieved by the classifiers underlines the separability and reliability of the identified clusters. Furthermore, feature importance analysis highlighted key genetic determinants within each cluster, offering actionable insights into potential biomarkers and critical genomic features. This framework bridges the gap between exploratory unsupervised learning and predictive supervised modeling, providing a scalable and interpretable method for analyzing complex genomic data sets. Its applicability extends to biomarker discovery, patient stratification, and other precision medicine applications, emphasizing its utility in advancing genomic research and clinical practice.

Humans↗

The potential of clustering methods for pre-test triage in sleep medicine: A systematic review.

Sleep disorders exhibit substantial heterogeneity, and traditional classifications may not fully capture clinically relevant subtypes. Clustering techniques can identify patient subgroups that improve phenotypic characterization and may support personalized management. This systematic review evaluated the application of clustering in sleep medicine, with particular focus on its potential use as a pre-test triage tool prior to formal sleep testing. PubMed/MEDLINE, Embase, Web of Science, and Scopus were searched to February 2025. Eligible studies applied clustering to classify sleep disorders in adults. Two reviewers independently conducted screening, data extraction, and risk-of-bias assessment using QUADAS-2. The protocol was registered on PROSPERO. Fifty-one studies (1983-2025) were included, predominantly focused on obstructive sleep apnea (OSA) (n = 38, 74%). Hierarchical clustering (n = 20) and K-means clustering (n = 14) were the most frequently used techniques. Internal validation was reported in only 18% of studies, and external validation was reported in only 1 study. Seven studies relied exclusively on baseline clinical, demographic, or questionnaire data, representing pre-test scenarios, whereas most incorporated polysomnography-derived variables, limiting their applicability to early clinical stratification. Hierarchical clustering was the most commonly applied method; however, the overall lack of validation limits confidence in the robustness and clinical applicability of identified phenotypes. The potential role of clustering as a pre-test triage strategy remains largely unexplored, as most studies focused on post-diagnostic phenotyping and were affected by incorporation bias. Future research should prioritize pre-test clinical variables, rigorously validate internally and externally, and adopt standardized methodological and reporting practices to facilitate clinical translation.

Humans↗

Fast tree search for a triangular lattice model of protein folding.

Using a triangular lattice model to study the designability of protein folding, we overcame the parity problem of previous cubic lattice model and enumerated all the sequences and compact structures on a simple two-dimensional triangular lattice model of size 4+5+6+5+4. We used two types of amino acids, hydrophobic and polar, to make up the sequences, and achieved 2(23)+2(12) different sequences excluding the reverse symmetry sequences. The total string number of distinct compact structures was 219,093, excluding reflection symmetry in the self-avoiding path of length 24 triangular lattice model. Based on this model, we applied a fast search algorithm by constructing a cluster tree. The algorithm decreased the computation by computing the objective energy of non-leaf nodes. The parallel experiments proved that the fast tree search algorithm yielded an exponential speed-up in the model of size 4+5+6+5+4. Designability analysis was performed to understand the search result.

Algorithms↗

DNA microarrays identification of primary and secondary target genes regulated by p53.

The transcriptional program regulated by the tumor suppressor p53 was analysed using oligonucleotide microarrays. A human lung cancer cell line that expresses the temperature sensitive murine p53 was utilized to quantitate mRNA levels of various genes at different time points after shifting the temperature to 32 degrees C. Inhibition of protein synthesis by cycloheximide (CHX) was used to distinguish between primary and secondary target genes regulated by p53. In the absence of CHX, 259 and 125 genes were up or down-regulated respectively; only 38 and 24 of these genes were up and down-regulated by p53 also in the presence of CHX and are considered primary targets in this cell line. Cluster analysis of these data using the super paramagnetic clustering (SPC) algorithm demonstrate that the primary genes can be distinguished as a single cluster among a large pool of p53 regulated genes. This procedure identified additional genes that co-cluster with the primary targets and can also be classified as such genes. In addition to cell cycle (e.g. p21, TGF-beta, Cyclin E) and apoptosis (e.g. Fas, Bak, IAP) related genes, the primary targets of p53 include genes involved in many aspects of cell function, including cell adhesion (e.g. Thymosin, Smoothelin), signaling (e.g. H-Ras, Diacylglycerol kinase), transcription (e.g. ATF3, LISCH7), neuronal growth (e.g. Ninjurin, NSCL2) and DNA repair (e.g. BTG2, DDB2). The results suggest that p53 activates concerted opposing signals and exerts its effect through a diverse network of transcriptional changes that collectively alter the cell phenotype in response to stress.

Animals↗

Protein subcellular localization prediction for Gram-negative bacteria using amino acid subalphabets and a combination of multiple support vector machines.

BACKGROUND: Predicting the subcellular localization of proteins is important for determining the function of proteins. Previous works focused on predicting protein localization in Gram-negative bacteria obtained good results. However, these methods had relatively low accuracies for the localization of extracellular proteins. This paper studies ways to improve the accuracy for predicting extracellular localization in Gram-negative bacteria. RESULTS: We have developed a system for predicting the subcellular localization of proteins for Gram-negative bacteria based on amino acid subalphabets and a combination of multiple support vector machines. The recall of the extracellular site and overall recall of our predictor reach 86.0% and 89.8%, respectively, in 5-fold cross-validation. To the best of our knowledge, these are the most accurate results for predicting subcellular localization in Gram-negative bacteria. CONCLUSION: Clustering 20 amino acids into a few groups by the proposed greedy algorithm provides a new way to extract features from protein sequences to cover more adjacent amino acids and hence reduce the dimensionality of the input vector of protein features. It was observed that a good amino acid grouping leads to an increase in prediction performance. Furthermore, a proper choice of a subset of complementary support vector machines constructed by different features of proteins maximizes the prediction accuracy.

Algorithms↗

Clustering and averaging of images in single-particle analysis.

Single particle analysis is a straightforward method for studying the structures of macromolecules that cannot be crystallized. It builds three-dimensional structures of particles by estimating the projection angles of their randomly oriented electron-microscopic images. The existing methods divide the images into clusters, build class averages for the clusters, and estimate the projection angle of each cluster. However, the clustering and the averaged images are highly sensitive to the choice of reference images and mask patterns for each cluster. Thus, the analyses are neither robust nor automatic, and their results depend heavily on the intuition and experience of researchers who set references. We have been developing a software system for single-particle analysis with new clustering and averaging algorithms for building the three-dimensional structures of target molecules. In this paper, we focus on the algorithms for the robust image-processing of the electron microscopic images in our system.

Algorithms↗

Advances in high-speed, three-dimensional imaging and automated segmentation algorithms for thick and overlapped clusters in cytologic preparations. Application to cervical smears.

OBJECTIVE: To use three-dimensional (3-D) imaging and localized adaptive image analysis to enable automated cervical smear screening systems to efficiently and effectively process thick and overlapped cell clusters currently left unprocessed. STUDY DESIGN: Instrumentation was developed to perform high-speed (50-200 optical sections per second at 256 x 256 resolution), 3-D imaging of thick regions of cervical smears. Normal and abnormal ThinPrep smears were imaged at two levels of resolution to approximate higher-resolution, wide-area imaging. Improved dual-resolution, 3-D image analysis algorithms were developed for segmenting nuclei in these clusters. RESULTS: Despite low contrast, high variability and dense overlaps, the algorithms detected 89% and correctly segmented 76% of nuclei in clusters from normal smears and detected 75% and correctly segmented 45% of nuclei in clusters from abnormal smears in low-resolution images. In high-resolution images they detected 88% and segmented 76% of nuclei from normal specimens and detected 55% and segmented 45% of nuclei from abnormal specimens. At least one nucleus from each cell cluster was correctly segmented. CONCLUSION: Selective application of 3-D imaging and 3-D image analysis to thick and overlapped regions can enable a significant fraction (45-89%) of clustered and embedded cells to be accessed by an automated analysis system. These regions are, for the most part, unprocessable by current two-dimensional methods.

Algorithms↗

Spatial genetic pattern in the land mollusc Helix aspersa inferred from a 'centre-based clustering' procedure.

The present work provides the first broad-scale screening of allozymes in the land snail Helix aspersa. By using overall information available on the distribution of genetic variation between 102 populations previously investigated, we expect to strengthen our knowledge on the spread of the invasive aspersa subspecies in the Western Mediterranean. We propose a new approach based on a centre-based clustering procedure to cluster populations into groups following rules of geographical proximity and genetic similarity. Assuming a stepping-stone model of diffusion, we apply a partitioning algorithm which clusters only populations that are geographically contiguous. The algorithm used, which is actually part of leading methods developed for analysing large microarray datasets, is that of the k-means. Its goal is to minimize the within-group variance. The spatial constraint is provided by a list of connections between localities deduced from a Delaunay network. After testing each optimal group for the presence of spatial arrangement in the genetic data, the inferred genetic structure was compared with partitions obtained from other methods published for defining homogeneous groups (i.e. the Monmonier and SAMOVA algorithms). Competing biogeographical scenarios inferred from the k-means procedure were then compared and discussed to shed more light on colonization routes taken by the species.

Animals↗

Class discovery analysis of the lung cancer gene expression data.

Traditional histological classification of lung cancer subtypes is informative, but incomplete. Recent studies of gene expression suggest that molecular classification can be used for effective diagnostic and prediction of the treatment outcome. We attempt to build a molecular classification based on the public data available from a few independent sources. The data is reanalyzed with a new cluster analysis algorithm. This algorithm allows us to preserve the high dimensionality of data and produce the cluster structure without preliminary selection of significant genes or any other presumption about the relation between different cancer and normal tissue samples. The resulting clusters are generally consistent with the histological classification. However, our analysis reveals many additional details and subtypes of previously defined types of lung cancer. Large histological cancer types can be further divided into subclasses with different patterns of gene expression. These subtypes should be taken into account in diagnostics, drug testing, and treatment development for lung cancer patients.

Cluster Analysis↗

Using hexamers to predict cis-regulatory motifs in Drosophila.

BACKGROUND: Cis-regulatory modules (CRMs) are short stretches of DNA that help regulate gene expression in higher eukaryotes. They have been found up to 1 megabase away from the genes they regulate and can be located upstream, downstream, and even within their target genes. Due to the difficulty of finding CRMs using biological and computational techniques, even well-studied regulatory systems may contain CRMs that have not yet been discovered. RESULTS: We present a simple, efficient method (HexDiff) based only on hexamer frequencies of known CRMs and non-CRM sequence to predict novel CRMs in regulatory systems. On a data set of 16 gap and pair-rule genes containing 52 known CRMs, predictions made by HexDiff had a higher correlation with the known CRMs than several existing CRM prediction algorithms: Ahab, Cluster Buster, MSCAN, MCAST, and LWF. After combining the results of the different algorithms, 10 putative CRMs were identified and are strong candidates for future study. The hexamers used by HexDiff to distinguish between CRMs and non-CRM sequence were also analyzed and were shown to be enriched in regulatory elements. CONCLUSION: HexDiff provides an efficient and effective means for finding new CRMs based on known CRMs, rather than known binding sites.

Algorithms↗

Transcriptome analysis of zebrafish embryogenesis using microarrays.

Zebrafish (Danio rerio) is a well-recognized model for the study of vertebrate developmental genetics, yet at the same time little is known about the transcriptional events that underlie zebrafish embryogenesis. Here we have employed microarray analysis to study the temporal activity of developmentally regulated genes during zebrafish embryogenesis. Transcriptome analysis at 12 different embryonic time points covering five different developmental stages (maternal, blastula, gastrula, segmentation, and pharyngula) revealed a highly dynamic transcriptional profile. Hierarchical clustering, stage-specific clustering, and algorithms to detect onset and peak of gene expression revealed clearly demarcated transcript clusters with maximum gene activity at distinct developmental stages as well as co-regulated expression of gene groups involved in dedicated functions such as organogenesis. Our study also revealed a previously unidentified cohort of genes that are transcribed prior to the mid-blastula transition, a time point earlier than when the zygotic genome was traditionally thought to become active. Here we provide, for the first time to our knowledge, a comprehensive list of developmentally regulated zebrafish genes and their expression profiles during embryogenesis, including novel information on the temporal expression of several thousand previously uncharacterized genes. The expression data generated from this study are accessible to all interested scientists from our institute resource database (http://giscompute.gis.a-star.edu.sg/~govind/zebrafish/data_download.html).

Journal Article↗

MELDB: a database for microbial esterases and lipases.

MELDB is a comprehensive protein database of microbial esterases and lipases which are hydrolytic enzymes important in the modern industry. Proteins in MELDB are clustered into groups according to their sequence similarities based on a local pairwise alignment algorithm and a graph clustering algorithm (TribeMCL). This differs from traditional approaches that use global pairwise alignment and joining methods. Our procedure was able to reduce the noise caused by dubious alignment in the distantly related or unrelated regions in the sequences. In the database, 883 esterase and lipase sequences derived from microbial sources are deposited and conserved parts of each protein are identified. HMM profiles of each cluster were generated to classify unknown sequences. Contents of the database can be keyword-searched and query sequences can be aligned to sequence profiles and sequences themselves.

Amino Acid Sequence↗

A graph-based clustering method for a large set of sequences using a graph partitioning algorithm.

A graph-based clustering method is proposed to cluster protein sequences into families, which automatically improves clusters of the conventional single linkage clustering method. Our approach formulates sequence clustering problem as a kind of graph partitioning problem in a weighted linkage graph, which vertices correspond to sequences, edges correspond to higher similarities than given threshold and are weighted by their similarities. The effectiveness of our method is shown in comparison with InterPro families in all mouse proteins in SWISS-PROT. The result clusters match to InterPro families much better than the single linkage clustering method. 77% of proteins in InterPro families are classified into appropriate clusters.

Algorithms↗

Identification of co-regulated genes through Bayesian clustering of predicted regulatory binding sites.

The identification of co-regulated genes and their transcription-factor binding sites (TFBS) are key steps toward understanding transcription regulation. In addition to effective laboratory assays, various computational approaches for the detection of TFBS in promoter regions of coexpressed genes have been developed. The availability of complete genome sequences combined with the likelihood that transcription factors and their cognate sites are often conserved during evolution has led to the development of phylogenetic footprinting. The modus operandi of this technique is to search for conserved motifs upstream of orthologous genes from closely related species. The method can identify hundreds of TFBS without prior knowledge of co-regulation or coexpression. Because many of these predicted sites are likely to be bound by the same transcription factor, motifs with similar patterns can be put into clusters so as to infer the sets of co-regulated genes, that is, the regulons. This strategy utilizes only genome sequence information and is complementary to and confirmative of gene expression data generated by microarray experiments. However, the limited data available to characterize individual binding patterns, the variation in motif alignment, motif width, and base conservation, and the lack of knowledge of the number and sizes of regulons make this inference problem difficult. We have developed a Gibbs sampling-based Bayesian motif clustering (BMC) algorithm to address these challenges. Tests on simulated data sets show that BMC produces many fewer errors than hierarchical and K-means clustering methods. The application of BMC to hundreds of predicted gamma-proteobacterial motifs correctly identified many experimentally reported regulons, inferred the existence of previously unreported members of these regulons, and suggested novel regulons.

Algorithms↗

Structural refinement of protein segments containing secondary structure elements: Local sampling, knowledge-based potentials, and clustering.

In this article, we present an iterative, modular optimization (IMO) protocol for the local structure refinement of protein segments containing secondary structure elements (SSEs). The protocol is based on three modules: a torsion-space local sampling algorithm, a knowledge-based potential, and a conformational clustering algorithm. Alternative methods are tested for each module in the protocol. For each segment, random initial conformations were constructed by perturbing the native dihedral angles of loops (and SSEs) of the segment to be refined while keeping the protein body fixed. Two refinement procedures based on molecular mechanics force fields - using either energy minimization or molecular dynamics - were also tested but were found to be less successful than the IMO protocol. We found that DFIRE is a particularly effective knowledge-based potential and that clustering algorithms that are biased by the DFIRE energies improve the overall results. Results were further improved by adding an energy minimization step to the conformations generated with the IMO procedure, suggesting that hybrid strategies that combine both knowledge-based and physical effective energy functions may prove to be particularly effective in future applications.

Algorithms↗

Clustered bottlenecks in mRNA translation and protein synthesis.

Using a model based on the totally asymmetric exclusion process, we investigate the effects of slow codons along messenger RNA. Ribosome density profiles near neighboring clusters of slow codons interact, enhancing suppression of ribosome throughput when such bottlenecks are closely spaced. Increasing the slow codon cluster size beyond approximately 3-4 codons does not significantly reduce the ribosome current. Our results are verified by both extensive Monte Carlo simulations and numerical calculation, and provide a biologically motivated explanation for the experimentally observed clustering of low-usage codons.

Algorithms↗