PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Ensemble attribute profile clustering: discovering and characterizing groups of genes with similar patterns of biological features.

BACKGROUND: Ensemble attribute profile clustering is a novel, text-based strategy for analyzing a user-defined list of genes and/or proteins. The strategy exploits annotation data present in gene-centered corpora and utilizes ideas from statistical information retrieval to discover and characterize properties shared by subsets of the list. The practical utility of this method is demonstrated by employing it in a retrospective study of two non-overlapping sets of genes defined by a published investigation as markers for normal human breast luminal epithelial cells and myoepithelial cells. RESULTS: Each genetic locus was characterized using a finite set of biological properties and represented as a vector of features indicating attributes associated with the locus (a gene attribute profile). In this study, the vector space models for a pre-defined list of genes were constructed from the Gene Ontology (GO) terms and the Conserved Domain Database (CDD) protein domain terms assigned to the loci by the gene-centered corpus LocusLink. This data set of GO- and CDD-based gene attribute profiles, vectors of binary random variables, was used to estimate multiple finite mixture models and each ensuing model utilized to partition the profiles into clusters. The resultant partitionings were combined using a unanimous voting scheme to produce consensus clusters, sets of profiles that co-occurred consistently in the same cluster. Attributes that were important in defining the genes assigned to a consensus cluster were identified. The clusters and their attributes were inspected to ascertain the GO and CDD terms most associated with subsets of genes and in conjunction with external knowledge such as chromosomal location, used to gain functional insights into human breast biology. The 52 luminal epithelial cell markers and 89 myoepithelial cell markers are disjoint sets of genes. Ensemble attribute profile clustering-based analysis indicated that both lists contained groups of genes with the functional properties of membrane receptor biology/signal transduction and nucleic acid binding/transcription. A subset of the luminal markers was associated with metabolic and oxidoreductase activities, whereas a subset of myoepithelial markers was associated with protein hydrolase activity. CONCLUSION: Given a set of genes and/or proteins associated with a phenomenon, process or system of interest, ensemble attribute profile clustering provides a simple method for collating and sythesizing the annotation data pertaining to them that are present in text-based, gene-centered corpora. The results provide information about properties common and unique to subsets of the list and hence insights into the biology of the problem under investigation.

Algorithms↗

Rapid assessment of extremal statistics for gapped local alignment.

The statistical significance of gapped local alignments is characterized by analyzing the extremal statistics of the scores obtained from the alignment of random amino acid sequences. By identifying a complete set of linked clusters, "islands," we devise a method which accurately predicts the extremal score statistics by using only one to a few pairwise alignments. The success of our method relies crucially on the link between the statistics of island scores and extremal score statistics. This link is motivated by heuristic arguments, and firmly established by extensive numerical simulations for a variety of scoring parameter settings and sequence lengths. Our approach is several orders of magnitude faster than the widely used shuffling method, since island counting is trivially incorporated into the basic Smith-Waterman alignment algorithm with minimal computational cost, and all islands are counted in a single alignment. The availability of a rapid and accurate significance estimation method gives one the flexibility to fine tune scoring parameters to detect weakly homologous sequences and obtain optimal alignment fidelity.

Algorithms↗

Analysis of a short test battery for children.

Explored the hierarchical structure of mental abilities by comparing principal component analysis with a hierarchical cluster analysis algorithm on a short test battery for children. This short test battery included the Wechsler Intelligence Scale for Children-Revised (WISC-R), the Peabody Individual Achievement Test (PIAT), the Beery Developmental Test of Visual-Motor Integration (VMI) and the Peabody Picture Vocabulary Test (PPVT). Scores from 182 children (116 boys and 66 girls) with a mean age of 10.83 were analyzed. The structure of the test battery included general intelligence, attention, academic achievement and perceptual-motor eye-hand coordination. Both the linear principal component analysis and the nonlinear hierarchical clustering analysis confirmed the hierarchical organization of mental abilities.

Achievement↗

[DNA chip data mining].

DNA chip data routinely contain gene expression levels of thousands of genes and the analysis should be supported by various computational tools. To be brief, the analysis procedure consists of four steps including image scanning, image processing, mathematical interpretation and biological interpretation. In image processing step, we should detect the spots and measure the signals of the spots and the background. In mathematical interpretation step, first of all we should massage the measured signals to make them appropriate for further mathematical analysis. The massaged data could be analyzed by various computational methods especially when the data were generated for multiple samples comparisons. The clustering techniques including hierarchical clustering, k-means clustering, SOTA, SOM are the most popular methods in this step. Various other multivariate statistics and related machine learning techniques are being introduced and applied to DNA chip data analysis recently. And finally the most important step we should tackle is the biological interpretation task. Although the depth of the domain knowledge about the biological situation under which the data were generated is the most important factor to elucidate the biological context, it could be supported by various bioinformatics tools including MEDLINE abstract processing by NLP techniques or genetic network models constructed by Boolean networks algorithms.

Algorithms↗

Parallel random tunneling algorithm for structural optimization of Lennard-Jones clusters up to N=330.

A random tunneling algorithm (RTA) is derived from the terminal repeller unconstrained subenergy tunneling (TRUST) algorithm, and the parallelization of the RTA is implemented with an island parallel paradigm. Combined with the techniques of angular moving, the parallel random tunneling algorithm (PRTA) is applied to the optimization of Lennard-Jones (LJ) atomic clusters, and all the global minima of LJ clusters with size up to 200 are successfully located. For the optimization of larger cluster, a PRTA with an improved seeding technique is developed and successfully applied to the optimization of LJ151-LJ309. Furthermore, the optimized structures of LJ309-330 with the PRTA, which have not been studied before, are also provided.

Journal Article↗

Relations between cognitive and symptom profile heterogeneity in schizophrenia.

Although numerous studies have consistently revealed cognitive heterogeneity in schizophrenia, the relationships between such heterogeneity and clinical phenomenology are not clear. Clusters derived from cognitive heterogeneity studies may or may not be associated with symptom profile or severity of illness. The purpose of this study was to examine the relationship between cognitive heterogeneity and demographic and clinical phenomenological measures. We examined cognitive heterogeneity in schizophrenia by empirically deriving clusters of patients based upon WAIS-R subtest scores and then analyzed the way in which these clusters related to demographic and symptom variables and to DSM-III-R diagnostic subtypes. Four cognitive clusters were identified that were consistent with previous research. These clusters were differentiated on the basis of educational level and occupational status but not on the basis of symptom profile, severity, or DSM-III-R subtypes. Results suggest that cognitive measures are independent of severity of the disorder and phenomenological symptom presentation in these subgroups of schizophrenic patients.

Adult↗

Clustering gene expression data with temporal abstractions.

This paper describes a new technique for clustering short time series coming from gene expression data. The technique is based on the labelling of the time series through temporal trend abstractions and a consequent clustering of the series on the basis of their labels. Clustering is performed at three different levels of aggregation of the original time series, so that the results are organized and visualized as a three-levels hierarchical tree. Results on simulated and on yeast data are shown. The technique appears robust and efficient and the results obtained are easy to be interpreted.

Algorithms↗

The role of long-range interactions in defining the secondary structure of proteins is overestimated.

MOTIVATION: Secondary structure predictions based on the properties of individual residues, and sometimes on local interactions, usually fail to exceed 65% efficiency. Therefore, non-local, long-range interactions seem to be a significant cause of this limitation. RESULTS: In this paper, we apply approaches to localize highly interacting residues and clusters of residues involved in multiple non-local interactions, and test various secondary structure predictions on this separate subset to assess the effect of long-range interactions on the prediction efficiencies. It was found that only a marginal part of the failure of secondary structure predictions results from the presence of long-range interactions. Alternative possibilities are also discussed.

Algorithms↗

Pvclust: an R package for assessing the uncertainty in hierarchical clustering.

SUMMARY: Pvclust is an add-on package for a statistical software R to assess the uncertainty in hierarchical cluster analysis. Pvclust can be used easily for general statistical problems, such as DNA microarray analysis, to perform the bootstrap analysis of clustering, which has been popular in phylogenetic analysis. Pvclust calculates probability values (p-values) for each cluster using bootstrap resampling techniques. Two types of p-values are available: approximately unbiased (AU) p-value and bootstrap probability (BP) value. Multiscale bootstrap resampling is used for the calculation of AU p-value, which has superiority in bias over BP value calculated by the ordinary bootstrap resampling. In addition the computation time can be enormously decreased with parallel computing option.

Algorithms↗

libcov: a C++ bioinformatic library to manipulate protein structures, sequence alignments and phylogeny.

BACKGROUND: An increasing number of bioinformatics methods are considering the phylogenetic relationships between biological sequences. Implementing new methodologies using the maximum likelihood phylogenetic framework can be a time consuming task. RESULTS: The bioinformatics library libcov is a collection of C++ classes that provides a high and low-level interface to maximum likelihood phylogenetics, sequence analysis and a data structure for structural biological methods. libcov can be used to compute likelihoods, search tree topologies, estimate site rates, cluster sequences, manipulate tree structures and compare phylogenies for a broad selection of applications. CONCLUSION: Using this library, it is possible to rapidly prototype applications that use the sophistication of phylogenetic likelihoods without getting involved in a major software engineering project. libcov is thus a potentially valuable building block to develop in-house methodologies in the field of protein phylogenetics.

Algorithms↗

Independent component analysis reveals new and biologically significant structures in micro array data.

BACKGROUND: An alternative to standard approaches to uncover biologically meaningful structures in micro array data is to treat the data as a blind source separation (BSS) problem. BSS attempts to separate a mixture of signals into their different sources and refers to the problem of recovering signals from several observed linear mixtures. In the context of micro array data, "sources" may correspond to specific cellular responses or to co-regulated genes. RESULTS: We applied independent component analysis (ICA) to three different microarray data sets; two tumor data sets and one time series experiment. To obtain reliable components we used iterated ICA to estimate component centrotypes. We found that many of the low ranking components indeed may show a strong biological coherence and hence be of biological significance. Generally ICA achieved a higher resolution when compared with results based on correlated expression and a larger number of gene clusters with significantly enriched for gene ontology (GO) categories. In addition, components characteristic for molecular subtypes and for tumors with specific chromosomal translocations were identified. ICA also identified more than one gene clusters significant for the same GO categories and hence disclosed a higher level of biological heterogeneity, even within coherent groups of genes. CONCLUSION: Although the ICA approach primarily detects hidden variables, these surfaced as highly correlated genes in time series data and in one instance in the tumor data. This further strengthens the biological relevance of latent variables detected by ICA.

Algorithms↗

A novel method for automated EMG decomposition and MUAP classification.

OBJECTIVE: This paper proposes a novel method for the extraction and classification of individual motor unit action potentials (MUAPs) from intramuscular electromyographic signals. METHODOLOGY: The proposed method automatically detects the number of template MUAP clusters and classifies them into normal, neuropathic or myopathic. It consists of three steps: (i) preprocessing of electromyogram (EMG) recordings, (ii) MUAP detection and clustering and (iii) MUAP classification. RESULTS: The approach has been validated using a dataset of EMG recordings and an annotated collection of MUAPs. The correct identification rate for MUAP clustering is 93, 95 and 92% for normal, myopathic and neuropathic, respectively. Ninety-one percent of the superimposed MUAPs were correctly identified. The obtained accuracy for MUAP classification is about 86%. CONCLUSION: The proposed method, apart from efficient EMG decomposition addresses automatic MUAP classification to neuropathic, myopathic or normal classes directly from raw EMG signals.

Action Potentials↗

Finding dominant sets in microarray data.

Clustering allows us to extract groups of genes that are tightly coexpressed from Microarray data. In this paper, a new method DSF_Clust is developed to find dominant sets (clusters). We have preformed DSF_Clust on several gene expression datasets and given the evaluation with some criteria. The results showed that this approach could cluster dominant sets of good quality compared to kmeans method. DSF_Clust deals with three issues that have bedeviled clustering, some dominant sets being statistically determined in a significance level, predefining cluster structure being not required, and the quality of a dominant set being ensured. We have also applied this approach to analyze published data of yeast cell cycle gene expression and found some biologically meaningful gene groups to be dug out. Furthermore, DSF_Clust is a potentially good tool to search for putative regulatory signals.

Algorithms↗

Automated de novo identification of repeat sequence families in sequenced genomes.

Repetitive sequences make up a major part of eukaryotic genomes. We have developed an approach for the de novo identification and classification of repeat sequence families that is based on extensions to the usual approach of single linkage clustering of local pairwise alignments between genomic sequences. Our extensions use multiple alignment information to define the boundaries of individual copies of the repeats and to distinguish homologous but distinct repeat element families. When tested on the human genome, our approach was able to properly identify and group known transposable elements. The program, should be useful for first-pass automatic classification of repeats in newly sequenced genomes.

Algorithms↗

ProteomeGRID: towards a high-throughput proteomics pipeline through opportunistic cluster image computing for two-dimensional gel electrophoresis.

The quest for high-throughput proteomics has revealed a number of critical issues. Whilst improved two-dimensional gel electrophoresis (2-DE) sample preparation, staining and imaging issues are being actively pursued by industry, reliable high-throughput spot matching and quantification remains a significant bottleneck in the bioinformatics pipeline, thus restricting the flow of data to mass spectrometry through robotic spot excision and protein digestion. To this end, it is important to establish a full multi-site Grid infrastructure for the processing, archival, standardisation and retrieval of proteomic data and metadata. Particular emphasis needs to be placed on large-scale image mining and statistical cross-validation for reliable, fully automated differential expression analysis, and the development of a statistical 2-DE object model and ontology that underpins the emerging HUPO PSI GPS (Human Proteome Organization Proteomics Standards Initiative General Proteomics Standards). The first step towards this goal is to overcome the computational and communications burden entailed by the image analysis of 2-DE gels with Grid enabled cluster computing. This paper presents the proTurbo framework as part of the ProteomeGRID, which utilises Condor cluster management combined with CORBA communications and JPEG-LS lossless image compression for task farming. A novel probabilistic eager scheduler has been developed to minimise make-span, where tasks are duplicated in response to the likelihood of the Condor machines' owners evicting them. A 60 gel experiment was pair-wise image registered (3540 tasks) on a 40 machine Linux cluster. Real-world performance and network overhead was gauged, and Poisson distributed worker evictions were simulated. Our results show a 4:1 lossless and 9:1 near lossless image compression ratio and so network overhead did not affect other users. With 40 workers a 32x speed-up was seen (80% resource efficiency), and the eager scheduler reduced the impact of evictions by 58%.

Algorithms↗

Gene clustering by latent semantic indexing of MEDLINE abstracts.

MOTIVATION: A major challenge in the interpretation of high-throughput genomic data is understanding the functional associations between genes. Previously, several approaches have been described to extract gene relationships from various biological databases using term-matching methods. However, more flexible automated methods are needed to identify functional relationships (both explicit and implicit) between genes from the biomedical literature. In this study, we explored the utility of Latent Semantic Indexing (LSI), a vector space model for information retrieval, to automatically identify conceptual gene relationships from titles and abstracts in MEDLINE citations. RESULTS: We found that LSI identified gene-to-gene and keyword-to-gene relationships with high average precision. In addition, LSI identified implicit gene relationships based on word usage patterns in the gene abstract documents. Finally, we demonstrate here that pairwise distances derived from the vector angles of gene abstract documents can be effectively used to functionally group genes by hierarchical clustering. Our results provide proof-of-principle that LSI is a robust automated method to elucidate both known (explicit) and unknown (implicit) gene relationships from the biomedical literature. These features make LSI particularly useful for the analysis of novel associations discovered in genomic experiments. AVAILABILITY: The 50-gene document collection used in this study can be interactively queried at http://shad.cs.utk.edu/sgo/sgo.html.

Abstracting and Indexing↗

Advances in computers and image processing with applications in nuclear medicine.

The continuing advances in hardware performance had made many previously computationally unattractive methods feasible, an example being iterative reconstruction in tomography, which is now routine. Dynamic SPECT can also be performed. However the aim of image processing is not just to produce pretty pictures, but to extract good clinical information. The methods also need to incorporate clinical knowledge and be defined using clinical constraints. In general data in nuclear medicine are n-D, often 3-D plus time. Data reduction for example by the extraction of physiological information, is important. Such data are in any case hard to visualise without compression, for example some kind of dimensionality reduction, going from n-D to a 2-D "functional" image. Both linear and non-linear operations can be considered. To extract physiological data, we need to fit models. Two classes of method are important: data driven and hypothesis driven. Examples of data driven methods are principal component analysis and factor analysis, where the model is derived form the data. Hypothesis driven methods are all implicitly or explicitly based on model fitting. A preliminary data driven step followed by an hypothesis driven approach could be called constrained statistical image analysis. Examples are shown as used in nuclear medicine and are being extended to MRI. Another important problem considered is that of multi-modality image registration and fusion. Although many methods exist, all based on the minimisation of an appropriate distance functions between 2 image data sets such as mutual information, additional constraints are required when the images are not so similar. Additional constraints can be imposed by means of cluster analysis of the n-dimensional feature space. In the analysis of such data, tests against reference data sets (atlases) are required, normally requiring warping the data sets in space, for example by the use of optic flow, or some kind of diffusion equation. Real time analysis of data during acquisition can lead to optimisation of acquisition procedures. Incorporation of such image analysis into a decision support system is desirable.

Algorithms↗

Combining gene annotations and gene expression data in model-based clustering: weighted method.

It has been increasingly recognized that incorporating prior knowledge into cluster analysis can result in more reliable and meaningful clusters. In contrast to the standard modelbased clustering with a global mixture model, which does not use any prior information, a stratified mixture model was recently proposed to incorporate gene functions or biological pathways as priors in model-based clustering of gene expression profiles: various gene functional groups form the strata in a stratified mixture model. Albeit useful, the stratified method may be less efficient than the global analysis if the strata are non-informative to clustering. We propose a weighted method that aims to strike a balance between a stratified analysis and a global analysis: it weights between the clustering results of the stratified analysis and that of the global analysis; the weight is determined by data. More generally, the weighted method can take advantage of the hierarchical structure of most existing gene functional annotation systems, such as MIPS and Gene Ontology (GO), and facilitate choosing appropriate gene functional groups as priors. We use simulated data and real data to demonstrate the feasibility and advantages of the proposed method.

Algorithms↗