PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Analysis of the innate and adaptive phases of allograft rejection by cluster analysis of transcriptional profiles.

Both clinical and experimental observations suggest that allograft rejection is a complex process with multiple components that are, at least partially, functionally redundant. Studies using graft recipients deficient in various genes including chemokines, cytokines, and other immune-associated genes frequently produce a phenotype of delayed, but not indefinitely prevented, rejection. Only a small subset of genetic deletions (for example, TCR(alpha) or beta, MHC I and II, B7-1 and B7-2, and recombinase-activating gene) permit permanent graft acceptance suggesting that rejection is orchestrated by a complex network of interrelated inflammatory and immune responses. To investigate this complex process, we have used oligonucleotide microarrays to generate quantitative mRNA expression profiles following transplantation. Patterns of gene expression were confirmed with real-time PCR data. Hierarchical clustering algorithms clearly differentiated the early and late phases of rejection. Self-organizing maps identified clusters of coordinately regulated genes. Genes up-regulated during the early phase included genes with prior biological functions associated with ischemia, injury, and Ag-independent innate immunity, whereas genes up-regulated in the late phase were enriched for genes associated with adaptive immunity.

Animals↗

Automatic recognition of hydrophobic clusters and their correlation with protein folding units.

A method is described to objectively identify hydrophobic clusters in proteins of known structure. Clusters are found by examining a protein for compact groupings of side chains. Compact clusters contain seven or more residues, have an average of 65% hydrophobic residues, and usually occur in protein interiors. Although smaller clusters contain only side-chain moieties, larger clusters enclose significant portions of the peptide backbone in regular secondary structure. These clusters agree well with hydrophobic regions assigned by more intuitive methods and many larger clusters correlate with protein domains. These results are in striking contrast with the clustering algorithm of J. Heringa and P. Argos (1991, J Mol Biol 220:151-171). That method finds that clusters located on a protein's surface are not especially hydrophobic and average only 3-4 residues in size. Hydrophobic clusters can be correlated with experimental evidence on early folding intermediates. This correlation is optimized when clusters with less than nine hydrophobic residues are removed from the data set. This suggests that hydrophobic clusters are important in the folding process only if they have enough hydrophobic residues.

Algorithms↗

Classification of Arabidopsis thaliana gene sequences: clustering of coding sequences into two groups according to codon usage improves gene prediction.

While genomic sequences are accumulating, finding the location of the genes remains a major issue that can be solved only for about a half of them by homology searches. Prediction methods are thus required, but unfortunately are not fully satisfying. Most prediction methods implicitly assume a unique model for genes. This is an oversimplification as demonstrated by the possibility to group coding sequences into several classes in Escherichia coli and other genomes. As no classification existed for Arabidopsis thaliana, we classified genes according to the statistical features of their coding sequences. A clustering algorithm using a codon usage model was developed and applied to coding sequences from A. thaliana, E. coli, and a mixture of both. By using it, Arabidopsis sequences were clustered into two classes. The CU1 and CU2 classes differed essentially by the choice of pyrimidine bases at the codon silent sites: CU2 genes often use C whereas CU1 genes prefer T. This classification discriminated the Arabidopsis genes according to their expressiveness, highly expressed genes being clustered in CU2 and genes expected to have a lower expression, such as the regulatory genes, in CU1. The algorithm separated the sequences of the Escherichia-Arabidopsis mixed data set into five classes according to the species, except for one class. This mixed class contained 89 % Arabidopsis genes from CU1 and 11 % E. coli genes, mostly horizontally transferred. Interestingly, most genes encoding organelle-targeted proteins, except the photosynthetic and photoassimilatory ones, were clustered in CU1. By tailoring the GeneMark CDS prediction algorithm to the observed coding sequence classes, its quality of prediction was greatly improved. Similar improvement can be expected with other prediction systems.

Algorithms↗

Application of fuzzy c-means segmentation technique for tissue differentiation in MR images of a hemorrhagic glioblastoma multiforme.

The application of a raw data-based, operator-independent MR segmentation technique to differentiate boundaries of tumor from edema or hemorrhage is demonstrated. A case of a glioblastoma multiforme with gross and histopathologic correlation is presented. The MR image data set was segmented into tissue classes based on three different MR weighted image parameters (T1-, proton density-, and T2-weighted) using unsupervised fuzzy c-means (FCM) clustering algorithm technique for pattern recognition. A radiological examination of the MR images and correlation with fuzzy clustering segmentations was performed. Results were confirmed by gross and histopathology which, to the best of our knowledge, reports the first application of this demanding approach. Based on the results of neuropathologic correlation, the application of FCM MR image segmentation to several MR images of a glioblastoma multiforme represents a viable technique for displaying diagnostically relevant tissue contrast information used in 3D volume reconstruction. With this technique, it is possible to generate segmentation images that display clinically important neuroanatomic and neuropathologic tissue contrast information from raw MR image data.

Adult↗

Brain white and gray matter anatomy of MRI segmentation based on tissue evaluation.

Different approaches to gray and white matter measurements in magnetic resonance imaging (MRI) have been studied. For clinical use, the estimated values must be reliable and accurate when, unfortunately, many techniques fail on these criteria in an unrestricted clinical environment. A recent method for tissue clusterization in MRI analysis has the advantage of great simplicity, and it takes the account of partial volume effects. In this study, we will evaluate the intensity of MR sequences known as T1-weighted images in an axial sliced section. Intensity group clustering algorithms are proposed to achieve further diagnosis for brain MRI, which has been hardly studied. Subjective study has been suggested to evaluate the clustering group intensity in order to obtain the best diagnosis as well as better detection for the suspected cases. This technique makes use of image tissue biases of intensity value pixels to provide 2 regions of interest as techniques. Moreover, the original mathematic solution could still be used with a specific set of modern sequences. There are many advantages to generalize the solution, which give far more scope for application and greater accuracy.

Algorithms↗

Improving the prediction of final infarct size in acute stroke with bolus delay-corrected perfusion MRI measures.

PURPOSE: To investigate whether bolus delay-corrected dynamic susceptibility contrast (DSC) perfusion MRI measures allowed a more accurate estimation of eventual infarct volume in 14 acute stroke patients using a predictive tissue classifier algorithm. MATERIALS AND METHODS: Tissue classification was performed using a expectation maximization and k-means clustering algorithm utilizing diffusion and T2 measures (diffusion-weighted imaging [DWI], apparent diffusion coefficient [ADC], and T2) combined with uncorrected perfusion measures cerebral blood flow ((CBF) and mean transit time [MTT]), bolus delay-corrected perfusion measures (cCBF and cMTT), and bolus delay-corrected perfusion indices (cCBF and cMTT with bolus delay). RESULTS: The mean similarity index (SI), a kappa-based correlation statistic reflecting the pixel-by-pixel classification agreement between predicted and 30-day T2 lesion volumes, were 0.55 +/- 0.19, 0.61 +/- 0.15 (P < 0.02) and 0.60 +/- 0.17 (P <0.03), respectively. Spearman's correlation coefficients, comparing predicted and final lesion volumes were 0.56 (P < 0.05), 0.70 (P < 0.01), and 0.84 (P < 0.001), respectively. We found a more significant correlation between predicted infarct volumes derived from bolus delay-corrected perfusion measures than from conventional perfusion measures when combined with diffusion measures and compared with final lesion volumes measured on 30-day T2 MRI scans. CONCLUSION: Bolus delay-corrected perfusion measures enable an improved prediction of infarct evolution and evaluation of the hemodynamic status of neuronal tissue in acute stroke.

Acute Disease↗

3-D image analysis of intra-cerebral brain hemorrhage from digitized CT films.

A new 3-D technique for the segmentation and quantification of human spontaneous intra-cerebral brain hemorrhage (ICH) is presented in this paper. The algorithm for ICH primary region segmentation uses the spatially weighted K-means histogram-based clustering algorithm. The ICH edema region segmentation algorithm employs an iterative morphological processing of the ICH brain data. A volume rendering technique is used for the effective 3-D visualization of ICH segmented regions. A computer program is developed for use in the human spontaneous ICH study involving a large number of patients. Experimental measurements and visualization results are presented which were computed on real ICH patient brain data.

Algorithms↗

Automatic definition of recurrent local structure motifs in proteins.

An automatic procedure for defining recurrent folding motifs in proteins of known structure is described. These motifs are formed by short polypeptide fragments of equal size containing between four and seven residues. The method applies a classical clustering algorithm that operates on distances between selected backbone atoms. In one application, we use it to cluster all protein fragments into only four structural classes. This classification is rough considering the observed diversity of local structures, but comparable in homogeneity to the four classes of secondary structure (alpha-helix, beta-strand, turn and coil). Yet, it discriminates between extended and curved coil and distinguishes beta-bulges from beta-strands. In a second application, the clustering procedure is combined with assignment of backbone dihedral angles to allowed regions in the Ramachandran map. This produces an exhaustive repertoire of highly homogeneous families of structural motifs that contains all the beta-hairpins, beta alpha- and alpha beta-loops previously defined by manual procedures, and new structural families of which two examples, a beta alpha-loop and an alpha-helix beginning, are analyzed in detail. The described automatic procedures should be useful in categorizing structure information in proteins, thereby increasing our ability to analyze relations between structure and sequence.

Amino Acid Sequence↗

Learning systems in biosignal analysis.

In biosignal analysis, the utility of artificial neural networks (ANN) in classifying electromyographic (EMG) data trained with the momentum back propagation algorithm has recently been demonstrated. In the current study, the self-organizing feature map algorithm, the genetics-based machine learning (GBML) paradigm, and the K-means nearest neighbour clustering algorithm are applied on the same set of data. The aim of this exercise is to show how these three paradigms can be used in practice, given that their diagnostic performance is problem- and parameter-dependent. A total of 720 macro EMG recordings were carried out from four groups, from seven normal, nine motor neuron disease, 14 Becker's muscular dystrophy, and six spinal muscular atrophy subjects, respectively. Twenty-three of the subjects were used for training and 13 for evaluating the various models. For each subject, the mean and the standard deviation of the parameters (i) amplitude, (ii) area, (iii) average power and (iv) duration were extracted. The feature vector was structured in two different ways for input to the models: an eight-input feature vector that consisted of both the mean and the standard deviation of the four parameters measured, and a four-input feature vector that included only the mean of the parameters. Also, due to the heterogenous nature of the spinal muscular atrophy group, three class models that excluded this group were investigated. In general, self-organizing feature map and GBML models resulted in comparable diagnostic performance of the order of 80-90% correct classifications (CCs) score for the evaluation set, whereas the K-means nearest neighbour algorithm models gave lower percentage CCs. Furthermore, for all three learning paradigms: better diagnostic performance was obtained for the three class models compared with the four class models; similar diagnostic performance was obtained for both the eight- and four-input feature vectors. Finally, it is claimed that the proposed methodology followed in this work can be applied for the development of diagnostic systems in the analysis of biosignals.

Algorithms↗

Mining the National Cancer Institute Anticancer Drug Discovery Database: cluster analysis of ellipticine analogs with p53-inverse and central nervous system-selective patterns of activity.

The United States National Cancer Institute conducts an anticancer drug discovery program in which approximately 10,000 compounds are screened every year in vitro against a panel of 60 human cancer cell lines from different organs. To date, approximately 62,000 compounds have been tested in the program, and a large amount of information on their activity patterns has been accumulated. For the current study, anticancer activity patterns of 112 ellipticine analogs were analyzed with the use of a hierarchical clustering algorithm. A dramatic coherence between molecular structures and their activity patterns could be seen from the cluster tree: the first subgroup (compounds 1-66) consisted principally of normal ellipticines, whereas the second subgroup (compounds 67-112) consisted principally of N2-alkyl-substituted ellipticiniums. Almost all apparent discrepancies in this clustering were explainable on the basis of chemical transformation to active forms under cell culture conditions. Correlations of activity with p53 status and selective activity against cells of central nervous system origin made this data set of special interest to us. The ellipticiniums, but not the ellipticines, were more potent on average against p53 mutant cells than against p53 wild-type ones (i.e., they seemed to be "p53-inverse") in this short term assay. This study strongly supports the hypothesis that "fingerprint" patterns of activity in the National Cancer Institute in vitro cell screening program encode incisive information on the mechanisms of action and other biological behaviors of tested compounds. Insights gained by mining the activity patterns could contribute to our understanding of anticancer drugs and the molecular pharmacology of cancer.

Antineoplastic Agents↗

Molecular marker profiles predict locoregional control of head and neck squamous cell carcinoma in a randomized trial of continuous hyperfractionated accelerated radiotherapy.

PURPOSE: Identification of factors that assist prediction of tumor response to radiotherapy may aid in refining treatment strategies and improving outcome. Possible association of molecular marker expression profiles with locoregional control of head and neck squamous cell carcinoma was investigated in a randomized trial of conventional versus continuous hyperfractionated accelerated radiotherapy (CHART). EXPERIMENTAL DESIGN: Tumor material was obtained from 402 patients. Immunohistochemistry was used to assess Ki-67, CD31, p53, Bcl-2, and cyclin D1 expression. A hierarchical clustering algorithm with a Bayesian information criterion was used to group tumors with similar marker expression; resulting expression profiles were then compared in terms of their difference in outcome after CHART and conventionally fractionated radiotherapy. RESULTS: Molecular marker profile was an independent prognostic factor for locoregional control. This was confirmed in multivariate analysis, including clinical variables such as tumor and nodal status, primary site, histological grade, age, and gender (P < 0.001 and P = 0.006 for local and nodal relapse, respectively). In particular, Bcl-2-positive tumors responded significantly better than average in both arms of the trial. Tumors negative for p53- and Bcl-2, with high and randomly patterned Ki-67 expression, responded worse than average with no benefit from CHART. Tumors with similarly negative p53 and Bcl-2, but low Ki-67 staining, with an organized pattern, benefit significantly from CHART schedule. CONCLUSIONS: This study demonstrates the potential of molecular profiles to predict radiotherapy response of head and neck squamous cell carcinoma and for treatment stratification. Distinct expression profiles correlate with three distinct clinical phenotypes, including good locoregional control, poor locoregional control, and an outcome strongly dependent upon fractionation schedule.

Algorithms↗

A novel approach to phylogenetic tree construction using stochastic optimization and clustering.

BACKGROUND: The problem of inferring the evolutionary history and constructing the phylogenetic tree with high performance has become one of the major problems in computational biology. RESULTS: A new phylogenetic tree construction method from a given set of objects (proteins, species, etc.) is presented. As an extension of ant colony optimization, this method proposes an adaptive phylogenetic clustering algorithm based on a digraph to find a tree structure that defines the ancestral relationships among the given objects. CONCLUSION: Our phylogenetic tree construction method is tested to compare its results with that of the genetic algorithm (GA). Experimental results show that our algorithm converges much faster and also achieves higher quality than GA.

Algorithms↗

Spontaneous nocturnal growth hormone secretion in anorexia nervosa.

In anorexia nervosa, serum GH levels are increased under basal conditions and respond abnormally to provocative stimuli. We report here, for the first time, an analysis of pulsatile GH secretion in these patients performed by Cluster algorithm. Seven anorectic and six normal weight, healthy women underwent serial blood sampling at 20-min intervals form 2030-0830 h for GH estimation. The total area under the curve (AUC; micrograms per L/min) was elevated 4-fold in anorectic patients compared to controls (4743.0 +/- 1520.09 vs. 1148.6 +/- 519.27; P < 0.01), largely due to an increase in the non-pulsatile fraction (3212.5 +/- 990.45 vs. 378.7 +/- 123.27; P < 0.01). Accordingly, the valley mean value was higher in anorectic than in control subjects (5.9 +/- 2.25 vs. 1.0 +/- 1.30 micrograms/L; P < 0.01). Furthermore, pulsatile AUC was also greater in anorectic patients (1530.4 +/- 654.72 vs. 769.8 +/- 404.02; P < 0.01) due to a significant increase in GH peak frequency (5.0 +/- 0.81 vs. 3.0 +/- 0.89; P < 0.01). No correlations were observed in these patients between body mass index and any of the parameters of spontaneous GH release, whereas a positive correlation was found between insulin-like growth factor I levels and pulsatile AUC (r2 = 0.583; P < 0.05), peak height (r2 = 0.743; P = 0.01), peak increment (r2 = 0.801; P < 0.01), and GH valley mean (r2 = 0.576; P < 0.05). In conclusion, it appears that the enhanced GH secretion in anorexia nervosa is the result of an increased frequency of secretory pulses superimposed on enhanced tonic GH secretion. Although this latter is consistent with a reduction of hypothalamic SRIH tone, the former may be accounted for by an increased number of GHRH discharges. Considering that in normal weight and obese subjects parameters of GH release are negatively correlated with adiposity indexes, the lack of such a negative correlation in our patients suggests that the enhancement of spontaneous GH release in anorectic patients is not merely the consequence of malnutrition-dependent impairment of insulin-like growth factor I production, but reflects a more complex hypothalamic dysregulation of GH release.

Adolescent↗

[Genetic divergence of Far Eastern dace species belonging to the genus Tribolodon (Pisces, Cyprinidae) and closely related taxa].

Based on a biochemical-genetic approach, heterozygosity and divergence of structural genes of 30 enzyme loci were analyzed in six dace species. In addition, intra- and interspecific divergence of gene expression was analyzed based on a sample of 12 to 15 loci. Mean heterozygosities per individual varied as follows: Tribolodon species, Hobs = 0.007 +/- 0.007 and Hexp = 0.007 +/- 0.007; T. ezoe, Hobs = 0.045 +/- 0.016 and Hexp = 0.067 +/- 0.029. Several variants of genetic distances were estimated. Standard Nei's distances (DN) varied from 0.145 to 0.284 in four dace species studied. As related to Tribolodon dace species, the following genetic distances were obtained for two members of other genera: Pseudaspius leptocephalus, DN = 0.269; Leuciscus waleckii, DN = 0.769. Based on the distance matrices, different clustering algorithms were realized. The main feature shared by different dendrograms was a separate position of the cluster joining Far-Eastern dace species, to which P. leptocephalus and L. waleckii are successively added. Among the species studied, the proportion of loci similar by expression (E) varied from 87 to 100%. The greatest difference was found between anadromous and nonanadromous ecotypes of T. hakonensis, E = 67%. The following conclusions can be made: (1) Four studied species of the genus Tribolodon are rather well genetically differentiated. Diagnostic loci are available. (2) A nominal dace species, T. species, should be considered the fourth isolated species of this genus, which is confirmed by its recent zoological acceptance of this species. (3) The origin and divergence of dace species belonging to the genus Tribolodon are relatively late (1 to 3 Myr ago) historical events. (4) Taxonomically, the genus Tribolodon belong to the tribe Pseudaspinini together with P. leptocephalus, which is confirmed by genetic data. (5) Data on heterozygosity and the divergence of structural and regulatory elements of genome, along with the proposed scheme of speciation types, suggest the following speciation modes for the species studied: for four species, adaptive divergence and for two species, genetic transformation.

Animals↗

Distinct gene expression patterns in a tamoxifen-sensitive human mammary carcinoma xenograft and its tamoxifen-resistant subline MaCa 3366/TAM.

The reasons why human mammary tumors become resistant to tamoxifen therapy are mainly unknown. Changes in gene expression may occur as cells acquire resistance to antiestrogens. We therefore undertook a comparative gene expression analysis of tamoxifen-sensitive and tamoxifen-resistant human breast cancer in vivo models using Affymetrix oligonucleotide arrays to analyze differential gene expression. Total RNAs from the tamoxifen-sensitive patient-derived mammary carcinoma xenograft MaCa 3366 and the tamoxifen-resistant model MaCa 3366/TAM were hybridized to Affymetrix HuGeneFL and to Hu95Av2 arrays. Pairwise comparisons and clustering algorithms were applied to identify differentially expressed genes and patterns of gene expression. As revealed by cluster analysis, the tamoxifen-sensitive and the tamoxifen-resistant breast carcinomas differed regarding their gene expression pattern. More than 100 transcripts are changed in abundance in MaCa 3366/TAM as compared with MaCa 3366. Among the genes that are differentially expressed in the tamoxifen-resistant tumors, there are several IFN-inducible and estrogen-responsive genes, and genes known to be involved in breast carcinogenesis. The genes neuronatin (NNAT) and bone marrow stem cell antigen 2 (BST2) were sharply up-regulated in MaCa 3366/TAM. The differential expression of four genes (NNAT, BST2, IGFBP5, and BCAS1) was confirmed by Taqman PCR. Our results provide the starting point for deriving markers for tamoxifen resistance by differential gene expression profiling in a human breast cancer model of acquired tamoxifen resistance. Finally, genes whose expression profiles are distinctly changed between the two xenograft lines will be further evaluated as potential targets for diagnostic or therapeutic approaches of tamoxifen-resistant breast cancer.

Animals↗

CLU: a new algorithm for EST clustering.

BACKGROUND: The continuous flow of EST data remains one of the richest sources for discoveries in modern biology. The first step in EST data mining is usually associated with EST clustering, the process of grouping of original fragments according to their annotation, similarity to known genomic DNA or each other. Clustered EST data, accumulated in databases such as UniGene, STACK and TIGR Gene Indices have proven to be crucial in research areas from gene discovery to regulation of gene expression. RESULTS: We have developed a new nucleotide sequence matching algorithm and its implementation for clustering EST sequences. The program is based on the original CLU match detection algorithm, which has improved performance over the widely used d2_cluster. The CLU algorithm automatically ignores low-complexity regions like poly-tracts and short tandem repeats. CONCLUSION: CLU represents a new generation of EST clustering algorithm with improved performance over current approaches. An early implementation can be applied in small and medium-size projects. The CLU program is available on an open source basis free of charge. It can be downloaded from http://compbio.pbrc.edu/pti.

Algorithms↗

A heuristic Bayesian method for segmenting DNA sequence alignments and detecting evidence for recombination and gene conversion.

We propose a heuristic approach to the detection of evidence for recombination and gene conversion in multiple DNA sequence alignments. The proposed method consists of two stages. In the first stage, a sliding window is moved along the DNA sequence alignment, and phylogenetic trees are sampled from the conditional posterior distribution with MCMC. To reduce the noise intrinsic to inference from the limited amount of data available in the typically short sliding window, a clustering algorithm based on the Robinson-Foulds distance is applied to the trees thus sampled, and the posterior distribution over tree clusters is obtained for each window position. While changes in this posterior distribution are indicative of recombination or gene conversion events, it is difficult to decide when such a change is statistically significant. This problem is addressed in the second stage of the proposed algorithm, where the distributions obtained in the first stage are post-processed with a Bayesian hidden Markov model (HMM). The emission states of the HMM are associated with posterior distributions over phylogenetic tree topology clusters. The hidden states of the HMM indicate putative recombinant segments. Inference is done in a Bayesian sense, sampling parameters from the posterior distribution with MCMC. Of particular interest is the determination of the number of hidden states as an indication of the number of putative recombinant regions. To this end, we apply reversible jump MCMC, and sample the number of hidden states from the respective posterior distribution.

Actins↗

Visual representation of cell subpopulation from flow cytometry data.

Flow cytometric systems are useful for protein identification and expression analysis, especially characterizing particular lineage or sublineage of cells. We clustered flow cytometry data of bone marrow cells into subpopulations using a clustering algorithm with its physical characteristics (cell size and cell granularity) and different molecular composition (cell reactivity with monoclonal antibodies). To display the cell subpopulations, we created a colored map according to the mean of 5 flow cytometry parameters based on a cluster. Such a map can reveal subpopulation properties that are not evident in the widely used scatter plot.

Algorithms↗