PubMed HealthSearch

PubMed · 41450215

MegaPlantTF: a machine learning framework for comprehensive identification and classification of plant transcription factors.

Abstract

MOTIVATION: Understanding the role of transcription factors (TFs) in plants is essential for the study of gene regulation and various biological processes. However, both TF detection and classification remain challenging due to the great diversity and complexity of these proteins. Conventional approaches, such as BLAST, often suffer from high computational complexity and limited performance on less common TF families. RESULTS: We introduce MegaPlantTF, the first comprehensive machine learning and deep learning framework for the prediction (TF versus non-TF) and classification (family-level) of plant TFs. Our method employs k-mer-based protein representations and a two-stage architecture combining a deep feed-forward neural network with a stacking ensemble classifier. To ensure robust performance assessment, we report micro-, macro-, and weighted-average performance metrics, providing a holistic evaluation of both frequent and underrepresented TF families. Additionally, we employ threshold-based evaluation to calibrate confidence in TF detection. The results show that MegaPlantTF achieves strong accuracy and precision, particularly with a k-mer size of 3 and a classification threshold of 0.5, and maintains stable performance even under stringent thresholds. In addition to the standard cross-validation tests, a use case study on Sorghum bicolor confirms that our method performs strongly in the genome-wide analysis, making it highly suitable for large-scale TF identification and classification tasks. MegaPlantTF represents a novel contribution by integrating k-mer encoding, binary family-specific classifiers, and a two-stage stacking ensemble into a unified, reproducible framework for large-scale plant TF identification and classification. AVAILABILITY AND IMPLEMENTATION: MegaPlantTF is freely accessible through a public web server available at https://bioinformatics.um6p.ma/MegaPlantTF. The complete source code, including pretrained models and example datasets, is available at https://github.com/Bioinformatics-UM6P/MegaPlantTF.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Genereux Akotenou, Asmaa H Hassan, Morad M Mokhtar, Achraf El Allali. 2026-01-02. MegaPlantTF: a machine learning framework for comprehensive identification and classification of plant transcription factors.. https://doi.org/10.1093/bioinformatics%2Fbtaf678

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A transcription factor regulatory atlas for activity inference and perturbation prediction.

Inferring transcription factor (TF) activity from transcriptomes and predicting transcriptome-wide responses to TF perturbations remain challenging, in part because available TF-mRNA resources often face a trade-off between precision and coverage and typically lack signed regulatory information. Here, we present TFActProfiler, a TF-mRNA resource and computational framework that learns signed, quantitative TF-mRNA regulatory coefficients by integrating heterogeneous prior evidence (ChIP-based, motif-based, and curated TF-mRNA annotations) with large-scale bulk and single-cell RNA-seq atlases. TFActProfiler contains 2 606 176 signed TF-mRNA interactions and improves TF activity inference in TF knockdown benchmarks relative to widely used regulon resources while retaining broad TF and target coverage. In addition, because the same learned regulatory coefficients can be used to model downstream transcriptional effects, TFActProfiler enables prediction of transcriptome-wide gene expression responses to TF knockdown without training on task-matched perturbation data. When perturbation datasets are available, TFActProfiler can be further refined to achieve performance comparable to state-of-the-art machine-learning baselines. By providing a direction-aware representation of TF-mRNA regulation for both activity inference and perturbation-response modeling, TFActProfiler supports systematic dissection of gene regulatory programs across diverse cellular contexts.

Transcription Factors

Motif-Cluster: Motif driven prioritization of transcription factor binding clusters.

Genome-wide analyses of transcription factor (TF) motif binding sites have largely emphasized individual high-affinity sites, while overlooking the regulatory importance of locally repetitive motif clusters. Such clusters, including combinations of weak and strong binding sites, can collectively enhance TF occupancy and regulatory activity. Here we present Motif-Cluster, an open-source framework for motif-driven prioritization and visualization of TF binding clusters using sequence information alone. Motif-Cluster integrates a density-based clustering strategy with flexible modeling of binding-site gaps and affinity signals, enabling the identification and ranking of candidate regulatory regions without requiring experimental binding data. Through simulations and multiple real-data analyses, we show that combining gap distributions with binding affinity effectively balances cluster size and signal strength while reducing noise from weak sites. Application to ZNF410 successfully recovers the previously characterized binding clusters in the CHD4 promoter, which are conserved between human and mouse. Additional case studies involving PHB1, TWIST1, and EGR1 further demonstrate the general applicability of the method across diverse transcription factors. Motif-Cluster also provides intuitive visualization and reproducible workflows to facilitate interpretation of spatially dense motif patterns. Overall, Motif-Cluster offers a robust and flexible approach for prioritizing transcription factor regulatory regions from genome-wide motif scans, enabling biological discovery and guiding experimental design, particularly in settings where direct genome-wide binding assays are unavailable.

Transcription Factors

Positional grammar of transcription factor binding partitions developmental and stress-response regulation in plants.

Understanding how transcription factor binding site (TFBS) position influences gene regulation remains a fundamental challenge in plants. Here, we integrate conserved multiDAP TFBS maps for 244 transcription factors (TFs) with single-nucleus chromatin accessibility, cell type-resolved gene expression, and hormone-response datasets across Brassicaceae species to determine how TFBS position relates to regulatory function. Although conserved TFBSs are enriched near transcription start sites (TSSs), TSS-proximal accessibility poorly predicts cell type-specific expression. Instead, cell type-specific expression correlates best with conserved TFBSs embedded in cell type-restricted chromatin, with TF family-specific distributions across distal promoters and introns. In contrast, TSS-proximal TFBSs in broadly accessible chromatin are associated with rapid transcriptional responses to abiotic and biotic stress hormones. Coding sequence TFBSs mark a distinct regulatory context in which the same DNA sequence encodes both amino acid sequence and TF motifs, including evidence that CDS-localized ABR1 binding may contribute to repression during hormone response. Finally, distal upstream regions contain conserved multi-family TF clusters with enhancer-like features overlapping rare cell type-specific accessible chromatin and enriched near genes controlling embryonic, meristematic, and hormone-dependent developmental patterning. Together, these results support a positional grammar in which TFBS position and chromatin context jointly partition developmental, stress-responsive, and repressive regulatory output in plants.

Transcription Factors