PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Image analysis and quantification of atherosclerosis using MRI.

This paper describes an image processing, pattern recognition, and computer graphics system for the noninvasive identification and evaluation of atherosclerosis using multidimensional Magnetic Resonance Imaging (MRI). Particular emphasis has been placed on the problem of developing a pattern recognition system for noninvasively identifying the different plaque classes involved in atherosclerosis using minimal a priori information. This pattern recognition technique involves an extension of the ISODATA clustering algorithm to include an information theoretic criterion (Consistent Akaike Information Criterion) to provide a measure of the fit of the cluster composition at a particular iteration to the actual data. A rapid 3-D display system is also described for the simultaneous display of multiple data classes resulting from the tissue identification process. This work demonstrates the feasibility of developing a "high information content" display which will aid in the diagnosis and analysis of the atherosclerotic disease process. Such capability will permit detailed and quantitative studies to assess the effectiveness of therapies, such as drug, exercise, and dietary regimens.

Algorithms↗

CAGNet: a structure-aware clustering-alternated graph network for cell-cell interaction inference in spatial transcriptomics.

MOTIVATION: Understanding cell-cell interactions (CCIs) in spatial transcriptomics is crucial for uncovering the spatial organization and functional heterogeneity of tissues. However, existing graph-based models typically rely on static clustering or fixed adjacency structures, which limits their ability to capture dynamic cellular relationships. RESULTS: We propose CAGNet, a two-stage framework for CCI inference from spatial transcriptomics data. In Stage 1, a Graph Attention Network encoder with joint feature and graph reconstruction learns structure-aware node embeddings from spatial gene expression profiles. In Stage 2, an alternating optimization mechanism iteratively updates cluster centers via KL-guided soft assignment and refines node embeddings through spatial graph reconstruction, establishing a closed-loop between representation learning and clustering. Experiments on three 10x Genomics Visium datasets demonstrate that CAGNet consistently outperforms six CCI inference baselines across ACC, AUC, AP, Precision, Recall, and F1. CAGNet also achieves the highest Adjusted Rand Index on all three datasets against six spatial domain identification methods, confirming that the learned embeddings capture biologically relevant spatial organization. Information-theoretic analysis further shows that CAGNet retains the highest mutual information between input features and learned embeddings among all compared methods. Ablation studies and 5-fold cross-validation confirm the contribution of each component and the reproducibility of the results. AVAILABILITY: The proposed method is implemented in the CAGNet package available at http://github.com/mahan1233333-maker/CAGNet .

Spatial Transcriptomics↗

Multispectral magnetic resonance images segmentation using fuzzy Hopfield neural network.

This paper demonstrates a fuzzy Hopfield neural network for segmenting multispectral MR brain images. The proposed approach is a new unsupervised 2-D Hopfield neural network based upon the fuzzy clustering technique. Its implementation consists of the combination of 2-D Hopfield neural network and fuzzy c-means clustering algorithm in order to make parallel implementation for segmenting multispectral MR brain images feasible. For generating feasible results, a fuzzy c-means clustering strategy is included in the Hopfield neural network to eliminate the need for finding weighting factors in the energy function which is formulated and based on a basic concept commonly used in pattern classification, called the 'within-class scatter matrix' principle. The suggested fuzzy c-means clustering strategy has also been proven to be convergent and to allow the network to learn more effectively than the conventional Hopfield neural network. The experimental results show that a near optimal solution can be obtained using the fuzzy Hopfield neural network based on the within-class scatter matrix.

Algorithms↗

Groups of histopathologic abnormalities in brains of very low birthweight infants.

The neuropathologic changes in brains of very premature infants are well recognized but relatively few studies have attempted to identify if specific neuropathologic features cluster together. These data could assist in determining pathogenetic mechanisms of immature brain injury. The goal of this study is to identify which, if any, combinations of histologic features occur together. We identified the presence or absence of 19 histologic features in the brains of 67 infants from a multicenter study of 1,665 prematurely born infants whose birthweight was 500-1,500 grams. We used clustering algorithms and factor analysis to group pathologic features that occurred together. Our results indicate that certain histopathologic features do cluster. For example, telencephalic white matter astrocytosis occurs in 2 groups: 1) associated with amphophilic globules, and, 2) in an uncorrelated group, associated with focal macrophage deposits and coagulative necroses. Parenchymal hemorrhage was not found to be associated with any telencephalic leukoencephalopathy, regardless of whether characterized by rarefaction, astrocytosis, focal coagulative necroses, or foci of macrophages in the white matter. Intraventricular hemorrhage and germinal matrix hemorrhage were not seen together more often than by chance expectation. Intraventricular hemorrhage was only marginally associated with parenchymal hemorrhage. Our data indicate that specific histopathologic features tend to preferentially cluster with each other in groups. This clustering may represent the manifestation of a common mechanism for each. These data should be valuable indicators for future research attempting to establish pathogenesis.

Algorithms↗

Review of MR image segmentation techniques using pattern recognition.

This paper has reviewed, with somewhat variable coverage, the nine MR image segmentation techniques itemized in Table II. A wide array of approaches have been discussed; each has its merits and drawbacks. We have also given pointers to other approaches not discussed in depth in this review. The methods reviewed fall roughly into four model groups: c-means, maximum likelihood, neural networks, and k-nearest neighbor rules. Both supervised and unsupervised schemes require human intervention to obtain clinically useful results in MR segmentation. Unsupervised techniques require somewhat less interaction on a per patient/image basis. Maximum likelihood techniques have had some success, but are very susceptible to the choice of training region, which may need to be chosen slice by slice for even one patient. Generally, techniques that must assume an underlying statistical distribution of the data (such as LML and UML) do not appear promising, since tissue regions of interest do not usually obey the distributional tendencies of probability density functions. The most promising supervised techniques reviewed seem to be FF/NN methods that allow hidden layers to be configured as examples are presented to the system. An example of a self-configuring network, FF/CC, was also discussed. The relatively simple k-nearest neighbor rule algorithms (hard and fuzzy) have also shown promise in the supervised category. Unsupervised techniques based upon fuzzy c-means clustering algorithms have also shown great promise in MR image segmentation. Several unsupervised connectionist techniques have recently been experimented with on MR images of the brain and have provided promising initial results. A pixel-intensity-based edge detection algorithm has recently been used to provide promising segmentations of the brain. This is also an unsupervised technique, older versions of which have been susceptible to oversegmenting the image because of the lack of clear boundaries between tissue types or finding uninteresting boundaries between slightly different types of the same tissue. To conclude, we offer some remarks about improving MR segmentation techniques. The better unsupervised techniques are too slow. Improving speed via parallelization and optimization will improve their competitiveness with, e.g., the k-nn rule, which is the fastest technique covered in this review. Another area for development is dynamic cluster validity. Unsupervised methods need better ways to specify and adjust c, the number of tissue classes found by the algorithm. Initialization is a third important area of research. Many of the schemes listed in Table II are sensitive to good initialization, both in terms of the parameters of the design, as well as operator selection of training data.(ABSTRACT TRUNCATED AT 400 WORDS)

Algorithms↗

Stokesian Dynamics Simulations of Ferromagnetic Colloidal Dispersions in a Simple Shear Flow.

We have investigated the behavior of clusters of ferromagnetic particles in a colloidal dispersion subjected to a simple shear flow. To do so, the Stokesian dynamics method has been used under the assumption that the effect of Brownian motion is negligible. For the case of no shear flow, the aggregate structures obtained by the Stokesian dynamics simulations agree well with Monte Carlo results qualitatively. We can, therefore, conclude that the Stokesian dynamics simulations can capture thick chainlike clusters without introducing a specific clustering algorithm, which is indispensable for Monte Carlo simulations. The behavior of the thick chainlike clusters in a simple shear flow is summarized as follows. The thick chainlike clusters decline in the shear flow direction as time advances. Since longer clusters experience larger shear forces, it is difficult for them to survive in such a situation. The thick chainlike clusters, therefore, dissociate into some short clusters. Such clusters are relatively stable in a shear flow, so that they do not decrease significantly any more. The viscosities have a strong relationship with the internal structures of the aggregates. The instantaneous viscosities, therefore, fluctuate significantly for the case of the thick chainlike clusters. Copyright 1998 Academic Press.

Journal Article↗

Hybrid systems for virtual screening: interest of fuzzy clustering applied to olfaction.

Kohonen neural networks, also known as Self Organizing Map (SOM), offer a useful 2D representation of the compound distribution inside a large chemical database. This distribution results from the compound organization in a molecular diversity hyperspace derived from a large set of molecular descriptors. Fuzzy techniques based on the "concept of partial truth" reveal to be also a valuable tool for the direct exploitation of chemical databases or SOM. In such cases a fuzzy clustering algorithm is used. In this paper, a complete hybrid system, combining SOM and fuzzy clustering, is applied. As example, a series of olfactory compounds was selected. The complexity of such information is that a same compound may exhibit different odors. It is shown how fuzzy logic helps to have a better understanding of the organization of the compounds. These hybrid systems, using simultaneously SOM and fuzzy clustering, are foreseen as powerful tools for "virtual pre-screening".

Fuzzy Logic↗

Use of cluster analysis technique for computerized recognition and prognosis of myocardial vulnerability.

Variability of factors exerting influence on arrhythmia origination leads to the appearance of polymodal distribution in sample space. Therefore, for the approximation of feature distribution, distribution mixture is used. To estimate the number of mixture components and to determine its parameters, cluster algorithm is used. The basic task of the algorithm is to identify the accumulation of vectors in the parallelepiped of their distribution. The accumulations of points are determined by testing statistical hypothesis of uniformity. On the basis of accumulations, clusters are formed and the parameters of normal mixtures of classes are estimated. Analysis of error matrix for recognition of mixture, enable to establish the parameters of the decision rule. The above algorithm was applied for recognition and prognosis of vulnerability of reentry and focal source in experiments on the right rabbit's atrium by using the electrophysiological parameters. We studied 30 cases of reentry, 36 cases of focal sources and 165-arrhythmia-free cases. As a result, we established the 7-class normal mixture which enabled a more effective (96.4%) recognition of the vulnerability types and 89.9% prognosis by features: increase in latency (theta), width of the interval of latency distribution, ratio theta/R, where R-refractory period.

Algorithms↗

Analysis of the innate and adaptive phases of allograft rejection by cluster analysis of transcriptional profiles.

Both clinical and experimental observations suggest that allograft rejection is a complex process with multiple components that are, at least partially, functionally redundant. Studies using graft recipients deficient in various genes including chemokines, cytokines, and other immune-associated genes frequently produce a phenotype of delayed, but not indefinitely prevented, rejection. Only a small subset of genetic deletions (for example, TCR(alpha) or beta, MHC I and II, B7-1 and B7-2, and recombinase-activating gene) permit permanent graft acceptance suggesting that rejection is orchestrated by a complex network of interrelated inflammatory and immune responses. To investigate this complex process, we have used oligonucleotide microarrays to generate quantitative mRNA expression profiles following transplantation. Patterns of gene expression were confirmed with real-time PCR data. Hierarchical clustering algorithms clearly differentiated the early and late phases of rejection. Self-organizing maps identified clusters of coordinately regulated genes. Genes up-regulated during the early phase included genes with prior biological functions associated with ischemia, injury, and Ag-independent innate immunity, whereas genes up-regulated in the late phase were enriched for genes associated with adaptive immunity.

Animals↗

Automatic recognition of hydrophobic clusters and their correlation with protein folding units.

A method is described to objectively identify hydrophobic clusters in proteins of known structure. Clusters are found by examining a protein for compact groupings of side chains. Compact clusters contain seven or more residues, have an average of 65% hydrophobic residues, and usually occur in protein interiors. Although smaller clusters contain only side-chain moieties, larger clusters enclose significant portions of the peptide backbone in regular secondary structure. These clusters agree well with hydrophobic regions assigned by more intuitive methods and many larger clusters correlate with protein domains. These results are in striking contrast with the clustering algorithm of J. Heringa and P. Argos (1991, J Mol Biol 220:151-171). That method finds that clusters located on a protein's surface are not especially hydrophobic and average only 3-4 residues in size. Hydrophobic clusters can be correlated with experimental evidence on early folding intermediates. This correlation is optimized when clusters with less than nine hydrophobic residues are removed from the data set. This suggests that hydrophobic clusters are important in the folding process only if they have enough hydrophobic residues.

Algorithms↗

Classification of Arabidopsis thaliana gene sequences: clustering of coding sequences into two groups according to codon usage improves gene prediction.

While genomic sequences are accumulating, finding the location of the genes remains a major issue that can be solved only for about a half of them by homology searches. Prediction methods are thus required, but unfortunately are not fully satisfying. Most prediction methods implicitly assume a unique model for genes. This is an oversimplification as demonstrated by the possibility to group coding sequences into several classes in Escherichia coli and other genomes. As no classification existed for Arabidopsis thaliana, we classified genes according to the statistical features of their coding sequences. A clustering algorithm using a codon usage model was developed and applied to coding sequences from A. thaliana, E. coli, and a mixture of both. By using it, Arabidopsis sequences were clustered into two classes. The CU1 and CU2 classes differed essentially by the choice of pyrimidine bases at the codon silent sites: CU2 genes often use C whereas CU1 genes prefer T. This classification discriminated the Arabidopsis genes according to their expressiveness, highly expressed genes being clustered in CU2 and genes expected to have a lower expression, such as the regulatory genes, in CU1. The algorithm separated the sequences of the Escherichia-Arabidopsis mixed data set into five classes according to the species, except for one class. This mixed class contained 89 % Arabidopsis genes from CU1 and 11 % E. coli genes, mostly horizontally transferred. Interestingly, most genes encoding organelle-targeted proteins, except the photosynthetic and photoassimilatory ones, were clustered in CU1. By tailoring the GeneMark CDS prediction algorithm to the observed coding sequence classes, its quality of prediction was greatly improved. Similar improvement can be expected with other prediction systems.

Algorithms↗

Application of fuzzy c-means segmentation technique for tissue differentiation in MR images of a hemorrhagic glioblastoma multiforme.

The application of a raw data-based, operator-independent MR segmentation technique to differentiate boundaries of tumor from edema or hemorrhage is demonstrated. A case of a glioblastoma multiforme with gross and histopathologic correlation is presented. The MR image data set was segmented into tissue classes based on three different MR weighted image parameters (T1-, proton density-, and T2-weighted) using unsupervised fuzzy c-means (FCM) clustering algorithm technique for pattern recognition. A radiological examination of the MR images and correlation with fuzzy clustering segmentations was performed. Results were confirmed by gross and histopathology which, to the best of our knowledge, reports the first application of this demanding approach. Based on the results of neuropathologic correlation, the application of FCM MR image segmentation to several MR images of a glioblastoma multiforme represents a viable technique for displaying diagnostically relevant tissue contrast information used in 3D volume reconstruction. With this technique, it is possible to generate segmentation images that display clinically important neuroanatomic and neuropathologic tissue contrast information from raw MR image data.

Adult↗

3-D image analysis of intra-cerebral brain hemorrhage from digitized CT films.

A new 3-D technique for the segmentation and quantification of human spontaneous intra-cerebral brain hemorrhage (ICH) is presented in this paper. The algorithm for ICH primary region segmentation uses the spatially weighted K-means histogram-based clustering algorithm. The ICH edema region segmentation algorithm employs an iterative morphological processing of the ICH brain data. A volume rendering technique is used for the effective 3-D visualization of ICH segmented regions. A computer program is developed for use in the human spontaneous ICH study involving a large number of patients. Experimental measurements and visualization results are presented which were computed on real ICH patient brain data.

Algorithms↗

Automatic definition of recurrent local structure motifs in proteins.

An automatic procedure for defining recurrent folding motifs in proteins of known structure is described. These motifs are formed by short polypeptide fragments of equal size containing between four and seven residues. The method applies a classical clustering algorithm that operates on distances between selected backbone atoms. In one application, we use it to cluster all protein fragments into only four structural classes. This classification is rough considering the observed diversity of local structures, but comparable in homogeneity to the four classes of secondary structure (alpha-helix, beta-strand, turn and coil). Yet, it discriminates between extended and curved coil and distinguishes beta-bulges from beta-strands. In a second application, the clustering procedure is combined with assignment of backbone dihedral angles to allowed regions in the Ramachandran map. This produces an exhaustive repertoire of highly homogeneous families of structural motifs that contains all the beta-hairpins, beta alpha- and alpha beta-loops previously defined by manual procedures, and new structural families of which two examples, a beta alpha-loop and an alpha-helix beginning, are analyzed in detail. The described automatic procedures should be useful in categorizing structure information in proteins, thereby increasing our ability to analyze relations between structure and sequence.

Amino Acid Sequence↗

Learning systems in biosignal analysis.

In biosignal analysis, the utility of artificial neural networks (ANN) in classifying electromyographic (EMG) data trained with the momentum back propagation algorithm has recently been demonstrated. In the current study, the self-organizing feature map algorithm, the genetics-based machine learning (GBML) paradigm, and the K-means nearest neighbour clustering algorithm are applied on the same set of data. The aim of this exercise is to show how these three paradigms can be used in practice, given that their diagnostic performance is problem- and parameter-dependent. A total of 720 macro EMG recordings were carried out from four groups, from seven normal, nine motor neuron disease, 14 Becker's muscular dystrophy, and six spinal muscular atrophy subjects, respectively. Twenty-three of the subjects were used for training and 13 for evaluating the various models. For each subject, the mean and the standard deviation of the parameters (i) amplitude, (ii) area, (iii) average power and (iv) duration were extracted. The feature vector was structured in two different ways for input to the models: an eight-input feature vector that consisted of both the mean and the standard deviation of the four parameters measured, and a four-input feature vector that included only the mean of the parameters. Also, due to the heterogenous nature of the spinal muscular atrophy group, three class models that excluded this group were investigated. In general, self-organizing feature map and GBML models resulted in comparable diagnostic performance of the order of 80-90% correct classifications (CCs) score for the evaluation set, whereas the K-means nearest neighbour algorithm models gave lower percentage CCs. Furthermore, for all three learning paradigms: better diagnostic performance was obtained for the three class models compared with the four class models; similar diagnostic performance was obtained for both the eight- and four-input feature vectors. Finally, it is claimed that the proposed methodology followed in this work can be applied for the development of diagnostic systems in the analysis of biosignals.

Algorithms↗

Mining the National Cancer Institute Anticancer Drug Discovery Database: cluster analysis of ellipticine analogs with p53-inverse and central nervous system-selective patterns of activity.

The United States National Cancer Institute conducts an anticancer drug discovery program in which approximately 10,000 compounds are screened every year in vitro against a panel of 60 human cancer cell lines from different organs. To date, approximately 62,000 compounds have been tested in the program, and a large amount of information on their activity patterns has been accumulated. For the current study, anticancer activity patterns of 112 ellipticine analogs were analyzed with the use of a hierarchical clustering algorithm. A dramatic coherence between molecular structures and their activity patterns could be seen from the cluster tree: the first subgroup (compounds 1-66) consisted principally of normal ellipticines, whereas the second subgroup (compounds 67-112) consisted principally of N2-alkyl-substituted ellipticiniums. Almost all apparent discrepancies in this clustering were explainable on the basis of chemical transformation to active forms under cell culture conditions. Correlations of activity with p53 status and selective activity against cells of central nervous system origin made this data set of special interest to us. The ellipticiniums, but not the ellipticines, were more potent on average against p53 mutant cells than against p53 wild-type ones (i.e., they seemed to be "p53-inverse") in this short term assay. This study strongly supports the hypothesis that "fingerprint" patterns of activity in the National Cancer Institute in vitro cell screening program encode incisive information on the mechanisms of action and other biological behaviors of tested compounds. Insights gained by mining the activity patterns could contribute to our understanding of anticancer drugs and the molecular pharmacology of cancer.

Antineoplastic Agents↗

Spontaneous nocturnal growth hormone secretion in anorexia nervosa.

In anorexia nervosa, serum GH levels are increased under basal conditions and respond abnormally to provocative stimuli. We report here, for the first time, an analysis of pulsatile GH secretion in these patients performed by Cluster algorithm. Seven anorectic and six normal weight, healthy women underwent serial blood sampling at 20-min intervals form 2030-0830 h for GH estimation. The total area under the curve (AUC; micrograms per L/min) was elevated 4-fold in anorectic patients compared to controls (4743.0 +/- 1520.09 vs. 1148.6 +/- 519.27; P < 0.01), largely due to an increase in the non-pulsatile fraction (3212.5 +/- 990.45 vs. 378.7 +/- 123.27; P < 0.01). Accordingly, the valley mean value was higher in anorectic than in control subjects (5.9 +/- 2.25 vs. 1.0 +/- 1.30 micrograms/L; P < 0.01). Furthermore, pulsatile AUC was also greater in anorectic patients (1530.4 +/- 654.72 vs. 769.8 +/- 404.02; P < 0.01) due to a significant increase in GH peak frequency (5.0 +/- 0.81 vs. 3.0 +/- 0.89; P < 0.01). No correlations were observed in these patients between body mass index and any of the parameters of spontaneous GH release, whereas a positive correlation was found between insulin-like growth factor I levels and pulsatile AUC (r2 = 0.583; P < 0.05), peak height (r2 = 0.743; P = 0.01), peak increment (r2 = 0.801; P < 0.01), and GH valley mean (r2 = 0.576; P < 0.05). In conclusion, it appears that the enhanced GH secretion in anorexia nervosa is the result of an increased frequency of secretory pulses superimposed on enhanced tonic GH secretion. Although this latter is consistent with a reduction of hypothalamic SRIH tone, the former may be accounted for by an increased number of GHRH discharges. Considering that in normal weight and obese subjects parameters of GH release are negatively correlated with adiposity indexes, the lack of such a negative correlation in our patients suggests that the enhancement of spontaneous GH release in anorectic patients is not merely the consequence of malnutrition-dependent impairment of insulin-like growth factor I production, but reflects a more complex hypothalamic dysregulation of GH release.

Adolescent↗

[Genetic divergence of Far Eastern dace species belonging to the genus Tribolodon (Pisces, Cyprinidae) and closely related taxa].

Based on a biochemical-genetic approach, heterozygosity and divergence of structural genes of 30 enzyme loci were analyzed in six dace species. In addition, intra- and interspecific divergence of gene expression was analyzed based on a sample of 12 to 15 loci. Mean heterozygosities per individual varied as follows: Tribolodon species, Hobs = 0.007 +/- 0.007 and Hexp = 0.007 +/- 0.007; T. ezoe, Hobs = 0.045 +/- 0.016 and Hexp = 0.067 +/- 0.029. Several variants of genetic distances were estimated. Standard Nei's distances (DN) varied from 0.145 to 0.284 in four dace species studied. As related to Tribolodon dace species, the following genetic distances were obtained for two members of other genera: Pseudaspius leptocephalus, DN = 0.269; Leuciscus waleckii, DN = 0.769. Based on the distance matrices, different clustering algorithms were realized. The main feature shared by different dendrograms was a separate position of the cluster joining Far-Eastern dace species, to which P. leptocephalus and L. waleckii are successively added. Among the species studied, the proportion of loci similar by expression (E) varied from 87 to 100%. The greatest difference was found between anadromous and nonanadromous ecotypes of T. hakonensis, E = 67%. The following conclusions can be made: (1) Four studied species of the genus Tribolodon are rather well genetically differentiated. Diagnostic loci are available. (2) A nominal dace species, T. species, should be considered the fourth isolated species of this genus, which is confirmed by its recent zoological acceptance of this species. (3) The origin and divergence of dace species belonging to the genus Tribolodon are relatively late (1 to 3 Myr ago) historical events. (4) Taxonomically, the genus Tribolodon belong to the tribe Pseudaspinini together with P. leptocephalus, which is confirmed by genetic data. (5) Data on heterozygosity and the divergence of structural and regulatory elements of genome, along with the proposed scheme of speciation types, suggest the following speciation modes for the species studied: for four species, adaptive divergence and for two species, genetic transformation.

Animals↗