PubMed HealthSearch

Biomedical subjects

Jindong Li

Publications and source records attributed to Jindong Li.

2 recordsLinked to original sources

Natural variation in GmSOP5 regulates seed oil and protein content during soybean domestication.

Seed oil content, protein content, and yield are agronomically important, correlated traits that determine the economic value of soybean (Glycine max). However, improving seed quality and yield simultaneously is challenging because gains in one breeding target often compromise the other, and the genetic basis of this trade-off is poorly understood. Here, we performed a genome-wide association study of 429 diverse soybean accessions and identified Seed Oil and Protein 5 (SOP5), which encodes a kinesin protein, as a key locus associated with seed oil and protein content. Knockout and overexpression experiments demonstrated that GmSOP5 positively affects seed oil content and 100-seed weight and negatively influences seed protein content. GmSOP5 is located in a selective sweep region, and the domestication-related GmSOP5H1 allele is nearly fixed in cultivated soybean, contributing to increased seed size, weight, and oil content and reduced protein content. Field trials demonstrated that neither loss-of-function GmSOP5-edited mutants, which have increased seed protein content, nor GmSOP5-overexpression lines, which have increased seed oil content, differed significantly in yield from wild-type plants, because changes in plant architecture were offset by changes in seed weight. Our results shed light on soybean domestication and suggest how pleiotropy can be harnessed in breeding to enhance seed quality without compromising yield.

GWAS

NanoSSL: attention mechanism-based self-supervised learning method for protein identification using nanopores.

MOTIVATION: Nanopores are cutting-edge interdisciplinary tools that can analyze biomolecules at the single-molecule level for many applications, e.g. DNA sequencing. Efforts are underway to extend nanopores to proteomics, including the development of machine learning algorithms for protein sequencing and identification. However, single-molecule data are intrinsically noisy and hard to process. Moreover, the development and performance of machine learning for nanopore is jeopardized by data scarcity. Self-supervised learning is an emerging method that may yield advantages in nanopore scenarios. RESULTS: We propose and experimentally validate Nanopore analysis using Self-Supervised Learning (NanoSSL), a generative self-supervised learning framework based on attention mechanisms for the identification of protein signals from nanopores. Leveraging a two-step approach consisting of self-supervised pre-training and supervised fine-tuning, NanoSSL learns useful feature representations from empirical data to facilitate downstream classification tasks. Inspired by the concept of fragmentation in conventional protein sequencing technologies, during pretraining each translocation event is split into multiple non-overlapping fragments of equal size, some of which are randomly masked and reconstructed using a masked autoencoder. Learning the feature representations of the reconstructed nanopore events facilitates molecular identification in fine-tuning. In this study, we retested a publicly available nanopore multiplexed protein sensing dataset for model iteration, and subsequently measured Alzheimer's disease biomarker Aβ1-42 using homemade solid-state nanopores. Empirical results indicated NanoSSL achieved an unprecedented performance across four metrics: accuracy, precision, recall, and F1 score, when classifying two mutated Aβ1-42, E22G and G37R. The self-supervised learning and attention mechanism were verified as the source of performance gains. AVAILABILITY AND IMPLEMENTATION: The main program is available at https://doi.org/10.5281/zenodo.17172822.

Nanopores