PubMed Health⌕ Search

PubMed · 12715832

Clustering-based approaches to discovering and visualising microarray data patterns.

Abstract

This article focuses on clustering techniques for the analysis of microarray data and discusses contributions and applications for the implementation of intelligent diagnostic systems and therapy design studies. Approaches to validating and visualising expression clustering results and software and other relevant resources to support clustering-based analyses are reviewed. Finally, this paper addresses current limitations and problems that need to be investigated for the development of an advanced generation of pattern discovery tools.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Francisco Azuaje. 2003. Clustering-based approaches to discovering and visualising microarray data patterns.. https://doi.org/10.1093/bib%2F4.1.31

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Genome mining reveals an architecturally expanded pyoluteorin-associated biosynthetic gene cluster and a divergent flavin-dependent halogenase-like sequence in deep-sea Pseudomonas Aeruginosa from the Gulf of Guinea.

BACKGROUND: Marine deep-sea environments harbour microorganisms with extraordinary biosynthetic potential, yet their secondary metabolite repertoires remain largely uncharacterised. RESULTS: This study reports the isolation, phenotypic characterisation, and whole-genome analysis of Pseudomonas aeruginosa strain E1, recovered from deep Atlantic seawater (Gulf of Guinea, ~2500 m depth), which exhibits antifungal activity against multidrug-resistant Candida parapsilosis. Three presumptive P. aeruginosa isolates (E1, E17, and E44) showed > 99% 16S rRNA gene sequence identity to P. aeruginosa reference sequences, while whole-genome dDDH analysis of strain E1 yielded 95.2% (95% CI: 93.6-96.4%; formula d4) relative to the P. aeruginosa type strain DSM 50071ᵀ (= ATCC 10145ᵀ), supporting its species-level assignment. Antifungal screening and PCR-based detection of flavin-dependent halogenase genes identified strain E1 as the primary candidate for genomic investigation. Illumina whole-genome sequencing produced a 6.33 Mb draft genome assembly (113 contigs, 5862 protein-coding genes, 66.4% GC content). Genome mining with antiSMASH 8.0 identified 27 biosynthetic gene clusters (BGCs) spanning nonribosomal peptide synthetase (NRPS), polyketide synthase (PKS), phenazine, terpene, and metallophore pathways. Region 7.1 of strain E1 harbours a predicted 50.8 kb pyoluteorin-associated BGC, comprising 34 genes, substantially larger than its terrestrial counterpart (~ 22 kb, ~ 17 genes), and featuring nine transport genes and three regulatory elements. Phylogenetic analysis resolved three halogenase genes: ctg7_146 showed 98.7% amino acid identity to PltA, and ctg7_149 showed 99.2% amino acid identity to PltM, supporting their annotation as PltA-like and PltM-like components of the predicted pyoluteorin biosynthetic pathway. Among the characterised reference enzymes included in this analysis, ctg7_143 showed the highest amino acid identity to PltM from P. fluorescens Pf-5. However, the identity remained low at approximately 30.4%, supporting its placement as a divergent FDH-like sequence rather than a close PltM orthologue. CONCLUSION: This study provides the first comprehensive genomic characterisation of a pyoluteorin-BGC-harbouring marine P. aeruginosa strain, demonstrating conservation of the core biosynthetic machinery alongside an expanded transport architecture and a divergent FDH-like sequence that may represent a candidate for future biochemical investigation. These findings expand current knowledge of FDH-like sequence diversity in deep-sea bacteria and support further investigation of Gulf of Guinea microorganisms as a potential source of biosynthetic and enzymatic diversity.

Multigene Family↗

Discovery of antimicrobial peptides from incomplete biosynthetic gene clusters to combat multidrug-resistant bacteria.

The escalating crisis of multidrug-resistant bacteria necessitates innovative antibiotic discovery platforms. Conventional antimicrobial peptide (AMP) mining often relies on complete biosynthetic gene clusters (BGCs), leaving fragmented genomic resources underexplored. Here, we present an evolution-inspired approach to reconstruct and predict AMPs from partial BGCs. Applying this strategy to 954 Paenibacillus genomes identifies five polymyxin-like peptides, NP001-NP005, with broad in vitro activity. Crucially, in murine models of polymyxin-resistant infection, NP001 reduced bacterial burdens by up to 1,000-fold in a thigh infection model and improved survival (50% vs. 0%) in a lethal peritonitis model. Structural simulations and biophysical assays revealed that NP001 maintains high affinity for bacterial membranes and effectively binds to MCR-1-modified lipid A, a key colistin-resistance mechanism. Moreover, Leu at position 10 of NP001 plays a key role in antibacterial activity against MCR-1-resistant bacteria. Our work establishes a generalizable framework for AMP discovery and introduces a promising therapeutic candidate, NP001, which effectively counteracts polymyxin-resistant pathogens.

Multigene Family↗

Fine-grained structural classification of biosynthetic gene cluster-encoded products.

MOTIVATION: Biosynthetic gene clusters (BGCs) are responsible the biosynthesis of many natural products, including a multitude of effective therapeutics and their precursors. Advances in genomic data collection as well as computational techniques have made it possible to identify BGCs at scale. However, accurately determining the types of BGC-encoded products from genomic content remains elusive. RESULTS: Here, we introduce BGC annotation tool (BGCat), a machine learning method for fine-grained structural classification of BGC-encoded products, leveraging the NPClassifier natural product nomenclature. Our method leverages a pre-trained protein language model for creating meaningful gene representations and a deep neural network for class label prediction. We show the method outperforms state-of-the-art approaches in coarse-grained product classification and is effective for detailed classification. We implement a clustering-based augmentation strategy for BGC-product relationships, addressing a crucial gap in the available datasets. We then introduce the concept of product class profiles of gene cluster families (GCFs), associating each GCF with a probabilistic distribution of product types and offering a new perspective on GCF functions. Lastly, we use BGCat to provide new product class labels for over 100k BGCs in antiSMASH DB that presently have minimal information about their products. AVAILABILITY AND IMPLEMENTATION: The source code and trained model weights are freely available at https://github.com/HassounLab/BGCat.

Multigene Family↗