PubMed Health⌕ Search

PubMed · 15831790

Masking repeats while clustering ESTs.

Abstract

A problem in EST clustering is the presence of repeat sequences. To avoid false matches, repeats have to be masked. This can be a time-consuming process, and it depends on available repeat libraries. We present a fast and effective method that aims to eliminate the problems repeats cause in the process of clustering. Unlike traditional methods, repeats are inferred directly from the EST data, we do not rely on any external library of known repeats. This makes the method especially suitable for analysing the ESTs from organisms without good repeat libraries. We demonstrate that the result is very similar to performing standard repeat masking before clustering.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Korbinian Schneeberger, Ketil Malde, Eivind Coward, Inge Jonassen. 2005-04-14. Masking repeats while clustering ESTs.. https://doi.org/10.1093/nar%2Fgki511

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging data.

MOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD.

Cluster Analysis↗

Focused library design in GPCR projects on the example of 5-HT(2c) agonists: comparison of structure-based virtual screening with ligand-based search methods.

The aim of this study was to investigate the usefulness of structure-based virtual screening (VS) for focused library design in G protein-coupled receptors (GPCR) projects on the example of 5-HT(2c) agonists. We compared the performance of structure-based VS against two different homology models using FRED for docking and ScreenScore, FlexX, and PMF for rescoring with the results of 12 ligand-based similarity searches using four different query compounds and three different similarity metrics (Daylight, FTree, Phacir). The result of the similarity search showed much variation, from an enrichment factor up to 3.2 to worse than random, whereas the structure-based VS gave a more stable result with a constant enrichment factor around 2. Additionally, actives retrieved by the structure-based approach were more diverse than the actives among the top scorers of the similarity searches. Based on these results, we suggest basing a focused library design for a GPCR project on a combination of a ligand-based similarity search and structure-based docking.

Cluster Analysis↗