PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Clustering Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Identifying splits with clear separation: a new class discovery method for gene expression data.

We present a new class discovery method for microarray gene expression data. Based on a collection of gene expression profiles from different tissue samples, the method searches for binary class distinctions in the set of samples that show clear separation in the expression levels of specific subsets of genes. Several mutually independent class distinctions may be found, which is difficult to obtain from most commonly used clustering algorithms. Each class distinction can be biologically interpreted in terms of its supporting genes. The mathematical characterization of the favored class distinctions is based on statistical concepts. By analyzing three data sets from cancer gene expression studies, we demonstrate that our method is able to detect biologically relevant structures, for example cancer subtypes, in an unsupervised fashion.

Algorithms↗

OrthoGUI: graphical presentation of Orthostrapper results.

SUMMARY: Orthostrapper is a program that calculates orthology support values for pairs of sequences in a multiple alignment (Storm and Sonnhammer, Bioinformatics, 18, 92-99, 2002). Here we present OrthoGUI, a web interface and display tool for Orthostrapper analysis. OrthoGUI visualizes the Orthostrapper output in both tabular and tree representations, and can also apply a clustering algorithm to identify groups of multiple orthologs, which are indicated by colour coding. AVAILABILITY: http://www.cgb.ki.se/OrthoGUI CONTACT: erik.sonnhammer@cgb.ki.se

ATP-Binding Cassette Transporters↗

stDyer-image improves clustering analysis of spatially resolved transcriptomics and proteomics with morphological images.

MOTIVATION: Spatially resolved transcriptomics (SRT) and spatially resolved proteomics (SRP) data enable the study of gene expression and protein abundances within their precise spatial and cellular contexts in tissues. Certain SRT and SRP technologies also capture corresponding morphology images, adding another layer of valuable information. However, few existing methods developed for SRT data effectively leverage these supplementary images to enhance clustering performance. RESULTS: Here, we introduce stDyer-image, an end-to-end deep learning framework designed for clustering for SRT and SRP datasets with images. Unlike existing methods that utilize images to complement gene expression data, stDyer-image directly links image features to cluster labels. This approach draws inspiration from pathologists, who can visually identify specific cell types or tumor regions from morphological images without relying on gene expression or protein abundances. Benchmarks against state-of-the-art tools demonstrate that stDyer-image achieves superior performance in clustering. Moreover, it is capable of handling large-scale datasets across diverse technologies, making it a versatile and powerful tool for spatial omics analysis. AVAILABILITY AND IMPLEMENTATION: The source code of stDyer-image and detailed tutorials are available at https://github.com/ericcombiolab/stDyer-image.

Proteomics↗

Client-server environment for high-performance gene expression data analysis.

SUMMARY: We have developed a platform independent, flexible and scalable Java environment for high-performance large-scale gene expression data analysis, which integrates various computational intensive hierarchical and non-hierarchical clustering algorithms. The environment includes a powerful client for data preparation and results visualization, an application server for computation and an additional administration tool. The package is available free of charge for academic and non-profit institutions.

Computing Methodologies↗

Gene-Expression Omnibus integration and clustering tools in SeqExpress.

UNLABELLED: SeqExpress, a gene-expression analysis suite, has been extended to offer a number of cluster generation, refinement and visualization techniques. The cluster generation methods have been specialized to deal with aspects of the sparseness and extreme values that occur within microarray data. The results of such cluster analysis can then be refined using either: a functional enrichment based procedure, which examines each cluster to see if it possesses an unusually high or low concentration of ontology terms; or by using Expectation-Maximization to find a mixture of model based distributions within the datasets. Visualizations are provided both to explore and compare the results of the cluster generation algorithms. In addition, a tool has been developed which integrates SeqExpress with the Gene-Expression Omnibus repository. The tool provides seamless access to the large number of experimental results in the repository, so that they can be visualized and analysed locally using SeqExpress. AVAILABILITY: SeqExpress is available as a 6 MB download from http://www.seqexpress.com and runs under Windows. A server-based version is available and is required for the GEO integration. SeqExpress is not affiliated with any academic institution, funding body or commercial organization and is free to use by all.

Algorithms↗

TICO: a tool for improving predictions of prokaryotic translation initiation sites.

UNLABELLED: We provide the tool 'TICO' (Translation Initiation site COrrection) for improving the results of conventional gene finders for prokaryotic genomes with regard to exact localization of the translation initiation site (TIS). At the current state TICO provides an interface for direct post processing of the predictions obtained from the widely used program GLIMMER. Our program is based on a clustering algorithm for completely unsupervised scoring of potential TIS locations. AVAILABILITY: Our tool can be freely accessed through a web interface at http://tico.gobics.de/ CONTACT: maike@gobics.de

Algorithms↗

Haplotype-based linkage disequilibrium mapping via direct data mining.

MOTIVATION: With the availability of large-scale, high-density single-nucleotide polymorphism markers and information on haplotype structures and frequencies, a great challenge is how to take advantage of haplotype information in the association mapping of complex diseases in case-control studies. RESULTS: We present a novel approach for association mapping based on directly mining haplotypes (i.e. phased genotype pairs) produced from case-control data or case-parent data via a density-based clustering algorithm, which can be applied to whole-genome screens as well as candidate-gene studies in small genomic regions. The method directly explores the sharing of haplotype segments in affected individuals that are rarely present in normal individuals. The measure of sharing between two haplotypes is defined by a new similarity metric that combines the length of the shared segments and the number of common alleles around any marker position of the haplotypes, which is robust against recent mutations/genotype errors and recombination events. The effectiveness of the approach is demonstrated by using both simulated datasets and real datasets. The results show that the algorithm is accurate for different population models and for different disease models, even for genes with small effects, and it outperforms some recently developed methods.

Algorithms↗

Accessing bioscience images from abstract sentences.

Images (e.g., figures) are important experimental results that are typically reported in bioscience full-text articles. Biologists need to access images to validate research facts and to formulate or to test novel research hypotheses. On the other hand, biologists live in an age of information explosion. As thousands of biomedical articles are published every day, systems that help biologists efficiently access images in literature would greatly facilitate biomedical research. We hypothesize that much of image content reported in a full-text article can be summarized by the sentences in the abstract of the article. In our study, more than one hundred biologists had tested this hypothesis and more than 40 biologists had evaluated a novel user-interface BioEx that allows biologists to access images directly from abstract sentences. Our results show that 87.8% biologists were in favor of BioEx over two other baseline user-interfaces. We further developed systems that explored hierarchical clustering algorithms to automatically identify abstract sentences that summarize the images. One of the systems achieves a precision of 100% that corresponds to a recall of 4.6%.

Abstracting and Indexing↗

A framework for gene expression analysis.

MOTIVATION: Global gene expression measurements as obtained, for example, in microarray experiments can provide important clues to the underlying transcriptional control mechanisms and network structure of a biological cell. In the absence of a detailed understanding of this gene regulation, current attempts at classification of expression data rely on clustering and pattern recognition techniques employing ad-hoc similarity criteria. To improve this situation, a better understanding of the expected relationships between expression profiles of genes associated by biological function is required. RESULTS: It is shown that perturbation expansions familiar from biological systems theory make precise predictions for the types of relationships to be expected for expression profiles of biologically associated genes, even if the underlying biological factors responsible for this association are not known. Classification criteria are derived, most of which are not usually employed in clustering algorithms. The approach is illustrated by using the AtGenExpress Arabidopsis thaliana developmental expression map.

Algorithms↗

Empirical evaluation of genetic clustering methods using multilocus genotypes from 20 chicken breeds.

We tested the utility of genetic cluster analysis in ascertaining population structure of a large data set for which population structure was previously known. Each of 600 individuals representing 20 distinct chicken breeds was genotyped for 27 microsatellite loci, and individual multilocus genotypes were used to infer genetic clusters. Individuals from each breed were inferred to belong mostly to the same cluster. The clustering success rate, measuring the fraction of individuals that were properly inferred to belong to their correct breeds, was consistently approximately 98%. When markers of highest expected heterozygosity were used, genotypes that included at least 8-10 highly variable markers from among the 27 markers genotyped also achieved >95% clustering success. When 12-15 highly variable markers and only 15-20 of the 30 individuals per breed were used, clustering success was at least 90%. We suggest that in species for which population structure is of interest, databases of multilocus genotypes at highly variable markers should be compiled. These genotypes could then be used as training samples for genetic cluster analysis and to facilitate assignments of individuals of unknown origin to populations. The clustering algorithm has potential applications in defining the within-species genetic units that are useful in problems of conservation.

Algorithms↗

A focused microarray approach to functional glycomics: transcriptional regulation of the glycome.

Glycosylation is the most common posttranslational modification of proteins, yet genes relevant to the synthesis of glycan structures and function are incompletely represented and poorly annotated on the commercially available arrays. To fill the need for expression analysis of such genes, we employed the Affymetrix technology to develop a focused and highly annotated glycogene-chip representing human and murine glycogenes, including glycosyltransferases, nucleotide sugar transporters, glycosidases, proteoglycans, and glycan-binding proteins. In this report, the array has been used to generate glycogene-expression profiles of nine murine tissues. Global analysis with a hierarchical clustering algorithm reveals that expression profiles in immune tissues (thymus [THY], spleen [SPL], lymph node, and bone marrow [BM]) are more closely related, relative to those of nonimmune tissues (kidney [KID], liver [LIV], brain [BRN], and testes [TES]). Of the biosynthetic enzymes, those responsible for synthesis of the core regions of N- and O-linked oligosaccharides are ubiquitously expressed, whereas glycosyltransferases that elaborate terminal structures are expressed in a highly tissue-specific manner, accounting for tissue and ultimately cell-type-specific glycosylation. Comparison of gene expression profiles with matrix-assisted laser desorption ionization-time of flight (MALDI-TOF) profiling of N-linked oligosaccharides suggested that the alpha1-3 fucosyltransferase 9, Fut9, is the enzyme responsible for terminal fucosylation in KID and BRN, a finding validated by analysis of Fut9 knockout mice. Two families of glycan-binding proteins, C-type lectins and Siglecs, are predominately expressed in the immune tissues, consistent with their emerging functions in both innate and acquired immunity. The glycogene chip reported in this study is available to the scientific community through the Consortium for Functional Glycomics (CFG) (http://www.functionalglycomics.org).

Animals↗

Gene expression profiling of primary breast carcinomas using arrays of candidate genes.

Breast cancer is characterized by an important histoclinical heterogeneity that currently hampers the selection of the most appropriate treatment for each case. This problem could be solved by the identification of new parameters that better predict the natural history of the disease and its sensitivity to treatment. A large-scale molecular characterization of breast cancer could help in this context. Using cDNA arrays, we studied the quantitative mRNA expression levels of 176 candidate genes in 34 primary breast carcinomas along three directions: comparison of tumor samples, correlations of molecular data with conventional histoclinical prognostic features and gene correlations. The study evidenced extensive heterogeneity of breast tumors at the transcriptional level. A hierarchical clustering algorithm identified two molecularly distinct subgroups of tumors characterized by a different clinical outcome after chemotherapy. This outcome could not have been predicted by the commonly used histoclinical parameters. No correlation was found with the age of patients, tumor size, histological type and grade. However, expression of genes was differential in tumors with lymph node metastasis and according to the estrogen receptor status; ERBB2 expression was strongly correlated with the lymph node status (P < 0.0001) and that of GATA3 with the presence of estrogen receptors (P < 0.001). Thus, our results identified new ways to group tumors according to outcome and new potential targets of carcinogenesis. They show that the systematic use of cDNA array testing holds great promise to improve the classification of breast cancer in terms of prognosis and chemosensitivity and to provide new potential therapeutic targets.

Adult↗

Episodic leptin release is independent of luteinizing hormone secretion.

Several studies suggest that leptin modulates hypothalamic-pituitary-gonadal axis functions. Leptin may stimulate release of gonadotrophin releasing hormone (GnRH) from the hypothalamus and of gonadotrophins from the pituitary. A synchronicity of luteinizing hormone (LH) and leptin pulses has been described in healthy women and in patients with polycystic ovarian syndrome, suggesting that leptin may modulate the episodic secretion of LH. However, it has not been established whether LH regulates the episodic secretion of leptin. To further examine LH-leptin interactions, we studied the episodic fluctuations of circulating LH and leptin in two patients with Kallmann's syndrome (KS) before and on day 7 of pulsatile GnRH administration, and compared these with those observed in the early follicular phase of 10 regularly menstruating women divided into two control groups according to the body mass index of each patient. To assess episodic hormone secretion, blood samples were collected at 10 min intervals for 6 h, before and on day 7 of GnRH administration in KS patients, and during days 3-7 of the follicular phase in normally cycling women. LH and leptin concentrations were measured in all samples. For pulse analysis, the cluster algorithm was used. Before treatment, an apulsatile pattern with no endogenous LH pulsations was observed in both KS patients. However, leptin pulses were assessed in both women. During GnRH administration, pulsatile LH activity was achieved in both patients with pulse characteristics similar to those of the respective control group. Serum leptin concentrations and leptin pulsatile patterns were not modified. These results suggest that circulating leptin is probably not modulated by pulsatile GnRH-LH secretion.

Adult↗

Are circulating leptin and luteinizing hormone synchronized in patients with polycystic ovary syndrome?

Animal and human studies suggest that leptin modulates hypothalamic-pituitary-gonadal axis functions. Leptin may stimulate gonadotrophin-releasing hormone (GnRH) release from the hypothalamus and luteinizing hormone (LH) and follicle stimulating hormone (FSH) secretion from the pituitary. A synchronicity of LH and leptin pulses has been described in healthy women, suggesting that leptin probably also regulates the episodic secretion of LH. In some pathological conditions, such as polycystic ovarian syndrome (PCOS), LH-leptin interactions are not known. The aim of the present investigation was to assess the episodic fluctuations of circulating LH and leptin in PCOS patients compared to regularly menstruating women. Six PCOS patients and six normal cycling (NC) women of similar age and body mass index (BMI) were studied. To assess episodic hormone secretion, blood samples were collected at 10-min intervals for 6 h. LH and leptin concentrations were measured in all samples. For pulse analysis the cluster algorithm was used. To detect an interaction between LH and leptin pulses, an analysis of copulsatility was employed. LH concentrations were significantly higher in the PCOS group in comparison to NC women, however serum leptin concentrations and leptin pulse characteristics for PCOS patients did not differ from NC women. A strong synchronicity between LH and leptin pulses was observed in NC women; 11 coincident leptin pulses were counted with a phase shift of 0 min (P = 0.027), 18 pulses with a phase shift of -1 (P = 0.025) and 24 pulses with a phase shift of -2 (P = 0.028). PCOS patients also exhibited a synchronicity between LH and leptin pulses but weaker (only 20 of 39 pulses) and with a phase shift greater than in normal women, leptin pulses preceding LH pulses by 20 min (P = 0.0163). These results demonstrate that circulating leptin and LH are synchronized in normal women and patients with PCOS. The real significance of the apparent copulsatility between LH and leptin must be elucidated, as well as the mechanisms that account for the ultradian leptin release.

Adult↗

Secretory pattern of leptin and LH during lactational amenorrhoea in breastfeeding normal and polycystic ovarian syndrome women.

Several studies have suggested that leptin modulates hypothalamic-pituitary-gonadal axis function. A synchronicity of LH and leptin pulses has been described in healthy women and in patients with polycystic ovarian syndrome (PCOS), suggesting that leptin may modulate the episodic secretion of LH. The aim of the present investigation was to assess the episodic fluctuations of circulating LH and leptin during lactational amenorrhoea in fully breastfeeding normal and PCOS women at 4 and 8 weeks postpartum, in order to establish LH-leptin interactions in the reactivation of the gonadal axis during this period. Six lactating PCOS patients and six normal lactating women of similar age and body mass index were studied. During a 12 h period on the 4th and 8th weeks postpartum, blood samples were collected at 10 min intervals for 12 h (22:00-10:00). Serum LH and leptin concentrations were measured in all samples. For pulse analysis, the cluster algorithm was used. To detect an interaction between LH and leptin pulses, an analysis of co-pulsatility was employed. LH concentrations tended to increase in both groups between the 4th and 8th weeks postpartum; however, serum leptin concentrations were not modified. Leptin pulse frequencies were similar at the 4th and 8th weeks postpartum, and did not differ between groups. Moreover, leptin pulse frequency was higher than LH pulse frequency in both groups, and in the two study periods. There was no synchronicity between LH and leptin pulses, and there were no increments in leptin concentration during the night. The fact that leptin concentrations were not modified and no synchronicity between LH and leptin pulses was observed suggests that, during lactational amenorrhoea, circulating leptin is probably not involved as a primary signal in promoting the reactivation of pulsatile LH secretion.

Adult↗

Computer-assisted prediction, classification, and delimitation of protein binding sites in nucleic acids.

We present a method to determine the location and extent of protein binding regions in nucleic acids by computer-assisted analysis of sequence data. The program ConsIndex establishes a library of consensus descriptions based on sequence sets containing known regulatory elements. These defined consensus descriptions are used by the program ConsInspector to predict binding sites in new sequences. We show the programs to correctly determine the significant regions involved in transcriptional control of seven sequence elements. The internal profile of relative variability of individual nucleotide positions within these regions paralleled experimental profiles of biological significance. Consensus descriptions are determined by employing an anchored alignment scheme, the results of which are then evaluated by a novel method which is superior to cluster algorithms. The alignment procedure is able to include several closely related sequences without biasing the consensus description. Moreover, the algorithm detects additional elements on the basis of a moderate distance correlation and is capable of discriminating between real binding sites and false positive matches. The software is well suited to cope with the frequent phenomenon of optional elements present in a subset of functionally similar sequences, while taking maximal advantage of the existing sequence data base. Since it requires only a minimum of seven sequences for a single element, it is applicable to a wide range of binding sites.

Algorithms↗

WIT: integrated system for high-throughput genome sequence analysis and metabolic reconstruction.

The WIT (What Is There) (http://wit.mcs.anl.gov/WIT2/) system has been designed to support comparative analysis of sequenced genomes and to generate metabolic reconstructions based on chromosomal sequences and metabolic modules from the EMP/MPW family of databases. This system contains data derived from about 40 completed or nearly completed genomes. Sequence homologies, various ORF-clustering algorithms, relative gene positions on the chromosome and placement of gene products in metabolic pathways (metabolic reconstruction) can be used for the assignment of gene functions and for development of overviews of genomes within WIT. The integration of a large number of phylogenetically diverse genomes in WIT facilitates the understanding of the physiology of different organisms.

Databases, Factual↗

ACLAME: a CLAssification of Mobile genetic Elements.

The ACLAME database (http://aclame.ulb.ac.be) is a collection and classification of prokaryotic mobile genetic elements (MGEs) from various sources, comprising all known phage genomes, plasmids and transposons. In addition to providing information on the full genomes and genetic entities, it aims to build a comprehensive classification of the functional modules of MGEs at the protein, gene and higher levels. This first version contains a comprehensive classification of 5069 proteins from 119 DNA bacteriophages into over 400 functional families. This classification was produced automatically using TRIBE-MCL, a graph-theory-based Markov clustering algorithm that uses sequence measures as input, and then manually curated. Manual curation was aided by consulting annotations available in public databases retrieved through additional sequence similarity searches using Psi-Blast and Hidden Markov Models. The database is publicly accessible and open to expert volunteers willing to participate in its curation. Its web interface allows browsing as well as querying the classification. The main objectives are to collect and organize in a rational way the complexity inherent to MGEs, to extend and improve the inadequate annotation currently associated with MGEs and to screen known genomes for the validation and discovery of new MGEs.

Bacteriophages↗