PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

eBLOCKs: enumerating conserved protein blocks to achieve maximal sensitivity and specificity.

Classifying proteins into families and superfamilies allows identification of functionally important conserved domains. The motifs and scoring matrices derived from such conserved regions provide computational tools that recognize similar patterns in novel sequences, and thus enable the prediction of protein function for genomes. The eBLOCKs database enumerates a cascade of protein blocks with varied conservation levels for each functional domain. A biologically important region is most stringently conserved among a smaller family of highly similar proteins. The same region is often found in a larger group of more remotely related proteins with a reduced stringency. Through enumeration, highly specific signatures can be generated from blocks with more columns and fewer family members, while highly sensitive signatures can be derived from blocks with fewer columns and more members as in a superfamily. By applying PSI-BLAST and a modified K-means clustering algorithm, eBLOCKs automatically groups protein sequences according to different levels of similarity. Multiple sequence alignments are made and trimmed into a series of ungapped blocks. Motifs and position-specific scoring matrices were derived from eBLOCKs and made available for sequence search and annotation. The eBLOCKs database provides a tool for high-throughput genome annotation with maximal specificity and sensitivity. The eBLOCKs database is freely available on the World Wide Web at http://motif.stanford.edu/eblocks/ to all users for online usage. Academic and not-for-profit institutions wishing copies of the program may contact Douglas L. Brutlag (brutlag@stanford.edu). Commercial firms wishing copies of the program for internal installation may contact Jacqueline Tay at the Stanford Office of Technology Licensing (jacqueline.tay@stanford.edu; http://otl.stanford.edu/).

Algorithms↗

MulPSSM: a database of multiple position-specific scoring matrices of protein domain families.

Representation of multiple sequence alignments of protein families in terms of position-specific scoring matrices (PSSMs) is commonly used in the detection of remote homologues. A PSSM is generated with respect to one of the sequences involved in the multiple sequence alignment as a reference. We have shown recently that the use of multiple PSSMs corresponding to an alignment, with several sequences in the family used as reference, improves the sensitivity of the remote homology detection dramatically. MulPSSM contains PSSMs for a large number of sequence and structural families of protein domains with multiple PSSMs for every family. The approach involves use of a clustering algorithm to identify most distinct sequences corresponding to a family. With each one of the distinct sequences as reference, multiple PSSMs have been generated. The current release of MulPSSM contains approximately 33,000 and approximately 38,000 PSSMs corresponding to 7868 sequence and 2625 structural families. A RPS_BLAST interface allows sequence search against PSSMs of sequence or structural families or both. An analysis interface allows display and convenient navigation of alignments and domain hits. MulPSSM can be accessed at http://crick.mbu.iisc.ernet.in/~mulpssm.

Databases, Protein↗

TreeDet: a web server to explore sequence space.

The TreeDet (Tree Determinant) Server is the first release of a system designed to integrate results from methods that predict functional sites in protein families. These methods take into account the relation between sequence conservation and evolutionary importance. TreeDet fully analyses the space of protein sequences in either user-uploaded or automatically generated multiple sequence alignments. The methods implemented in the server represent three main classes of methods for the detection of family-dependent conserved positions, a tree-based method, a correlation based method and a method that employs a principal component analyses coupled to a cluster algorithm. An additional method is provided to highlight the reliability of the position in the alignments. The server is available at http://www.pdg.cnb.uam.es/servers/treedet.

Amino Acid Sequence↗

Mitochondrial DNA differentiation during the speciation process in Peromyscus.

We address the problem of the possible significance of biological speciation to the magnitude and pattern of divergence of asexually transmitted characters in bisexual species. The empirical data for this report consist of restriction endonuclease site variability in maternally transmitted mitochondrial DNA (mtDNA) isolated from 82 samples of Peromyscus polionotus and P. leucopus collected from major portions of the respective species' ranges. Data are analyzed together with previously published information on P. maniculatus, a sibling species to polionotus. Maps of restriction sites indicate that all of the variation observed can be reasonably attributed to base substitutions leading to loss or gain of particular recognition sites. Magnitude of mtDNA sequence divergence within polionotus (maximum approximately equal to 2%) is roughly comparable to that observed within any of five previously identified mtDNA assemblages in maniculatus. Sequence divergence within leucopus (maximum approximately equal to 4%) is somewhat greater than that within polionotus. Consideration of probable evolutionary links among mtDNA restriction site maps allowed estimation of matriarchal phylogenies within polionotus and leucopus. Clustering algorithms and qualitative Wagner procedures were used to generate phenograms and parsimony networks, respectively, for the between-species comparisons. Three simple graphical models are presented to illustrate some conceivable relationships of mtDNA differentiation to speciation. In theoretical case I, each of two reproductively defined species (A and B) is monophyletic in matriarchal genealogy; the common female ancestor of either species can either predate or postdate the speciation. In case II, neither species is monophyletic in matriarchal genotype. In case III, species B is monophyletic but forms a subclade within A which is thus paraphyletic with respect to B. The empirical results for mtDNA in maniculatus and polionotus appear to conform closely to case III. These theoretical and empirical considerations raise a number of questions about the general relationship of the speciation process to the evolution of uniparentally transmitted traits. Some of these considerations are presented, and it is suggested that the distribution patterns of mtDNA sequence variation within and among extant species should be of considerable relevance to the particular demographies of speciation.

Animals↗

Gradient correction and classification of CT lung images for the automated quantification of mosaic attenuation pattern.

PURPOSE: The detection of density differences, or "mosaic attenuation pattern," on CT images may be difficult when the regional inhomogeneity of the density of the lung parenchyma is subtle. The purpose of this work was to develop a fully automated method for the reproducible quantification of the underattenuated areas of the lung parenchyma. This technique may be useful in increasing the precision of investigation of structure/function relationships. METHOD: Anatomical segmentation was achieved by a structure-filtering operator based on mathematical morphology. To compensate for the density gradient visible on lung CT scans, a model-based iterative deconvolution filter and an adaptive clustering algorithm were developed. Validation was performed with CT images from a lung phantom, 15 patients with constrictive obliterative bronchiolitis, and 8 normal subjects. RESULTS: The accuracy of the estimate of the density gradient on phantom studies was 93.3%. The automated quantification of the areas of decreased attenuation on scans of constrictive obliterative bronchiolitis was within 8.2% from the average scoring of two experienced observers. CONCLUSION: The proposed technique is fully automated and can accurately correct for density gradient and classify areas of decreased attenuation on lung CT images.

Bronchiolitis Obliterans↗

Self-complexity and the persistence of depression.

Self-complexity, a measure of the structure of cognition involving the self, was used to predict the persistence of depression in patients diagnosed with major depression. Self-descriptions offered by depressed patients were analyzed using a clustering algorithm to model cognitive structure. Indices of positive and negative self-complexity, derived from the resulting models, were used to predict depressive symptomatology 9 months after the onset of a major depression. Negative self-complexity uniquely predicted subsequent levels of depression even after the effects of initial levels of depression, self-evaluation, and dysfunctional attitudes were statistically removed. Highly complex negative self-representation appears to be associated with poor recovery from a major depressive episode. Future studies examining the relationship between cognition and psychopathology should investigate, in addition to its content, the formal and structural properties of cognition.

Adult↗

Limited genetic exchanges between populations of an insect pest living on uncultivated and related cultivated host plants.

Habitats in agroecosystems are ephemeral, and are characterized by frequent disturbances forcing pest species to successively colonize various hosts belonging either to the cultivated or to the uncultivated part of the agricultural landscape. The role of wild habitats as reservoirs or refuges for the aphid Sitobion avenae that colonize cultivated fields was assessed by investigating the genetic structure of populations collected on both cereal crops (wheat, barley and oat) and uncultivated hosts (Yorkshire fog, cocksfoot, bulbous oatgrass and tall oatgrass) in western France. Classical genetic analyses and Bayesian clustering algorithms indicate that genetic differentiation is high between populations collected on uncultivated hosts and on crops, revealing a relatively limited gene flow between the uncultivated margins and the cultivated part of the agroecosystem. A closer genetic relatedness was observed between populations living on plants belonging to the same tribe (Triticeae, Poeae and Aveneae tribes) where aphid genotypes appeared not to be specialized on a single host, but rather using a group of related plant species. Causes of this ecological differentiation and its implications for integrated pest management of S. avenae as cereals pest are discussed.

Analysis of Variance↗

Profiles of amino-acid utilisation and production amongst strains of Branhamella catarrhalis.

Utilisation and production of amino acids by isolates of Branhamella catarrhalis was studied by ion exchange chromatography after cells had been grown in nutrient broth and Mueller-Hinton broth. The profiles of amino acids used and produced by each strain were compared by a single linkage cluster algorithm. The results of this study reflect the biochemical and physiological heterogeneity amongst strains of B. catarrhalis.

Amino Acids↗

Tissue gene expression analysis using arrayed normalized cDNA libraries.

We have used oligonucleotide-fingerprinting data on 60,000 cDNA clones from two different mouse embryonic stages to establish a normalized cDNA clone set. The normalized set of 5,376 clones represents different clusters and therefore, in almost all cases, different genes. The inserts of the cDNA clones were amplified by PCR and spotted on glass slides. The resulting arrays were hybridized with mRNA probes prepared from six different adult mouse tissues. Expression profiles were analyzed by hierarchical clustering techniques. We have chosen radioactive detection because it combines robustness with sensitivity and allows the comparison of multiple normalized experiments. Sensitive detection combined with highly effective clustering algorithms allowed the identification of tissue-specific expression profiles and the detection of genes specifically expressed in the tissues investigated. The obtained results are publicly available (http://www.rzpd.de) and can be used by other researchers as a digital expression reference.

Algorithms↗

A microarray-based antibiotic screen identifies a regulatory role for supercoiling in the osmotic stress response of Escherichia coli.

Changes in DNA supercoiling are induced by a wide range of environmental stresses in Escherichia coli, but the physiological significance of these responses remains unclear. We now demonstrate that an increase in negative supercoiling is necessary for transcriptional activation of a large subset of osmotic stress-response genes. Using a microarray-based approach, we have characterized supercoiling-dependent gene transcription by expression profiling under conditions of high salt, in conjunction with the microbial antibiotics novobiocin, pefloxacin, and chloramphenicol. Algorithmic clustering and statistical measures for gauging cellular function show that this subset is enriched for genes critical in osmoprotectant transport/synthesis and rpoS-driven stationary phase adaptation. Transcription factor binding site analysis also supports regulation by the global stress sigma factor rpoS. In addition, these studies implicate 60 uncharacterized genes in the osmotic stress regulon, and offer evidence for a broader role for supercoiling in the control of stress-induced transcription.

Anti-Bacterial Agents↗

Detection of fixed points in spatiotemporal signals by a clustering method.

We present a method to determine fixed points in spatiotemporal signals. The method combines a clustering algorithm and a nonlinear analysis method fitting temporal dynamics. A 144-dimensional simulated signal, similar to a Kueppers-Lortz instability, is analyzed and its fixed points are reconstructed.

Algorithms↗

Cluster monte carlo study of multicomponent fluids of the stillinger-helfand and widom-rowlinson type

Phase transitions of fluid mixtures of the type introduced by Stillinger and Helfand are studied using a continuum version of the invaded cluster algorithm. Particles of the same species do not interact, but particles of different types interact with each other via a repulsive potential. Examples of interactions include the Gaussian molecule potential and a repulsive step potential. Accurate values of the critical density, fugacity, and magnetic exponent are found in two and three dimensions for the two-species model. The effect of varying the number of species and of introducing quenched impurities is also investigated. In all the cases studied, mixtures of q species are found to have properties similar to q-state Potts models.

Journal Article↗

Monte Carlo study of the three-dimensional Coulomb frustrated Ising ferromagnet.

We have investigated, by Monte Carlo simulation, the phase diagram of a three-dimensional Ising model with nearest-neighbor ferromagnetic interactions and small, but long-range (Coulombic) antiferromagnetic interactions. We have developed an efficient cluster algorithm and used different lattice sizes and geometries, which allows us to obtain the main characteristics of the temperature-frustration phase diagram. Our finite-size scaling analysis confirms that the melting of the lamellar phases into the paramagnetic phase is driven first order by the fluctuations. Transitions between ordered phases with different modulation patterns are observed in some regions of the diagram, in agreement with a recent mean-field analysis.

Journal Article↗

Combination of improved multibondic method and the Wang-Landau method.

We propose a method for Monte Carlo simulation of statistical physical models with discretized energy. The method is based on several ideas including the cluster algorithm, the multicanonical Monte Carlo method and its acceleration proposed recently by Wang and Landau. As in the multibondic ensemble method proposed by Janke and Kappler, the present algorithm performs a random walk in the space of the bond population to yield the state density as a function of the bond number. A test on the Ising model shows that the number of Monte Carlo sweeps required of the present method for obtaining the density of state with a given accuracy is proportional to the system size, whereas it is proportional to the system size squared for other conventional methods. In addition, the method shows a better performance than the original Wang-Landau method in measurement of physical quantities.

Journal Article↗

Synchronization and coarsening (without self-organized criticality) in a forest-fire model.

We study the long-time dynamics of a forest-fire model with deterministic tree growth and instantaneous burning of entire forests by stochastic lightning strikes. Asymptotically the system organizes into a coarsening self-similar mosaic of synchronized patches within which trees regrow and burn simultaneously. We show that the average patch length grows linearly with time as t--> infinity. The number density of patches of length L, N(L,t), scales as -2N(L/ ), and within a mean-field rate equation description we find that this scaling function decays as N(x) approximately e(-1/x) for x-->0, and as e(-x) for x--> infinity. In one dimension, we develop an event-driven cluster algorithm to study the asymptotic behavior of large systems. Our numerical results are consistent with mean-field predictions for patch coarsening.

Journal Article↗

Monte Carlo simulation of a planar lattice model with P4 interactions.

Monte Carlo study of a two-dimensional lattice with three-dimensional spins (d=2,n=3) interacting with nearest neighbors via a -P4(cos theta) potential, where P4 is the fourth Legendre polynomial and theta is the angle between two spins, has been reported for lattice sizes ranging from 10 x 10 to 160 x 160. A cluster algorithm for spin updating with a histogram reweighting technique has been used and finite size scaling has been performed. The model exhibits a strong first order phase transition at a dimensionless temperature 0.376+/-0.015. The phase transition appears to be driven by condensation of topological defects and the defect density D increases sharply at the transition temperature. The temperature derivative dD/dT* is found to obey a linear scaling relation with the lattice size L. The behavior of the model seems to be remarkably different from the two-dimensional P2 model, that has been investigated by other authors, although both models possess the same symmetry and topological defects play an important role in the phase transition.

Journal Article↗

Three-dimensional mixed-wet random pore-scale network modeling of two- and three-phase flow in porous media. I. Model description.

We present a three-dimensional network model to simulate two- and three-phase capillary dominated processes at the pore level. The displacement mechanisms incorporated in the model are based on the physics of multiphase flow observed in micromodel experiments. All the important features of immiscible fluid flow at the pore scale, such as wetting layers, spreading layers of the intermediate-wet phase, hysteresis, and wettability alteration are implemented in the model. Wettability alteration allows any values for the advancing and receding oil-water, gas-water, and gas-oil contact angles to be assigned. Multiple phases can be present in each pore or throat (element), in wetting and spreading layers, as well as occupying the center of the pore space. In all, some 30 different generic fluid configurations for two- and three-phase flow are analyzed. Double displacement and layer formation are implemented as well as direct two-phase displacement and layer collapse events. Every element has a circular, square, or triangular cross section. A random network that represents the pore space in Berea sandstone is used in this study. The model computes relative permeabilities, saturation paths, and capillary pressures for any displacement sequence. A methodology to track a given three-phase saturation path is presented that enables us to compare predicted and measured relative permeabilities on a point-by-point basis. A robust displacement-based clustering algorithm is also presented.

Absorption↗

Bulk and surface phase transitions in the three-dimensional O4 spin model.

We investigate the O(4) spin model on the simple-cubic lattice by means of the Wolff cluster algorithm. Using the toroidal boundary condition, we locate the bulk critical point at coupling K(c) = 0.935 856(2), and determine the bulk thermal magnetic renormalization exponents as y(t) = 1.337 5(15) and y(h) = 2.482 0(2), respectively. The universal ratio Q=m(2)(2)/m(4) is also determined as 0.9142(1). The precision of these estimates significantly improves over that of the existing results. Then, we simulate the critical O(4) model with two open surfaces on which the coupling strength K(1) can be varied. At the ordinary transitions, the surface magnetic exponent is determined as y((o))(h1) = 1.020 2(12). Further, we find a so-called special surface transition at (k) = K(1)/K-1 = 1.258(20). At this point, the surface thermal exponent y(s)(t1) is rather close to zero, and we cannot exclude that the corresponding surface transition is Kosterlitz-Thouless-like. The surface magnetic exponent is y((s))/h1 = 1.816(2).

Journal Article↗