PubMed Health⌕ Search

Biomedical subjects

Yinghua Dong

Publications and source records attributed to Yinghua Dong.

4 recordsLinked to original sources

Analysis and functional annotation of expressed sequence tags from the fall armyworm Spodoptera frugiperda.

BACKGROUND: Little is known about the genome sequences of lepidopteran insects, although this group of insects has been studied extensively in the fields of endocrinology, development, immunity, and pathogen-host interactions. In addition, cell lines derived from Spodoptera frugiperda and other lepidopteran insects are routinely used for baculovirus foreign gene expression. This study reports the results of an expressed sequence tag (EST) sequencing project in cells from the lepidopteran insect S. frugiperda, the fall armyworm. RESULTS: We have constructed an EST database using two cDNA libraries from the S. frugiperda-derived cell line, SF-21. The database consists of 2,367 ESTs which were assembled into 244 contigs and 951 singlets for a total of 1,195 unique sequences. CONCLUSION: S. frugiperda is an agriculturally important pest insect and genomic information will be instrumental for establishing initial transcriptional profiling and gene function studies, and for obtaining information about genes manipulated during infections by insect pathogens such as baculoviruses.

Animals↗

Support vector machines in HTS data mining: Type I MetAPs inhibition study.

This article reports a successful application of support vector machines (SVMs) in mining high-throughput screening (HTS) data of a type I methionine aminopeptidases (MetAPs) inhibition study. A library with 43,736 small organic molecules was used in the study, and 1355 compounds in the library with 40% or higher inhibition activity were considered as active. The data set was randomly split into a training set and a test set (3:1 ratio). The authors were able to rank compounds in the test set using their decision values predicted by SVM models that were built on the training set. They defined a novel score PT50, the percentage of the test set needed to be screened to recover 50% of the actives, to measure the performance of the models. With carefully selected parameters, SVM models increased the hit rates significantly, and 50% of the active compounds could be recovered by screening just 7% of the test set. The authors found that the size of the training set played a significant role in the performance of the models. A training set with 10,000 member compounds is likely the minimum size required to build a model with reasonable predictive power.

Algorithms↗

Discover protein sequence signatures from protein-protein interaction data.

BACKGROUND: The development of high-throughput technologies such as yeast two-hybrid systems and mass spectrometry technologies has made it possible to generate large protein-protein interaction (PPI) datasets. Mining these datasets for underlying biological knowledge has, however, remained a challenge. RESULTS: A total of 3108 sequence signatures were found, each of which was shared by a set of guest proteins interacting with one of 944 host proteins in Saccharomyces cerevisiae genome. Approximately 94% of these sequence signatures matched entries in InterPro member databases. We identified 84 distinct sequence signatures from the remaining 172 unknown signatures. The signature sharing information was then applied in predicting sub-cellular localization of yeast proteins and the novel signatures were used in identifying possible interacting sites. CONCLUSION: We reported a method of PPI data mining that facilitated the discovery of novel sequence signatures using a large PPI dataset from S. cerevisiae genome as input. The fact that 94% of discovered signatures were known validated the ability of the approach to identify large numbers of signatures from PPI data. The significance of these discovered signatures was demonstrated by their application in predicting sub-cellular localizations and identifying potential interaction binding sites of yeast proteins.

Binding Sites↗

Tobacco genes induced by the bacterial effector protein AvrPto.

The type III effector protein AvrPto acts as a virulence factor in susceptible plants lacking a cognate resistance gene but triggers hypersensitive response and disease resistance in tomato plants carrying the Pto gene or in tobacco plants carrying an unknown resistance gene. To assist the characterization of cellular responses caused by AvrPto in the plant, a pathogen-free system was adopted to isolate genes up-regulated 12 h after induced expression of AvrPto. By using subtraction cloning and transgenic tobacco plants expressing avrPto as a transgene, we isolated 125 nonredundant cDNA clones that represent avrPto-response genes (ARG). In addition to genes that are known to be induced by Pto-avrPto recognition, a number of new genes were also isolated. Most of ARG showed a specific induction in tobacco plants challenged with incompatible or nonhost pathogens. The use of an avrPto mutant that selectively eliminated the avrPto recognition in tobacco demonstrated that the ARG were induced in a highly specific manner by the avirulence, instead of the virulence activity of avrPto.

Bacterial Proteins↗