PubMed Health⌕ Search

Biomedical subjects

Zhenkun Gou

Publications and source records attributed to Zhenkun Gou.

3 recordsLinked to original sources

Using evolutionary and structural information to predict DNA-binding sites on DNA-binding proteins.

Proteins that interact with DNA are involved in a number of fundamental biological activities such as DNA replication, transcription, and repair. A reliable identification of DNA-binding sites in DNA-binding proteins is important for functional annotation, site-directed mutagenesis, and modeling protein-DNA interactions. We apply Support Vector Machine (SVM), a supervised pattern recognition method, to predict DNA-binding sites in DNA-binding proteins using the following features: amino acid sequence, profile of evolutionary conservation of sequence positions, and low-resolution structural information. We use a rigorous statistical approach to study the performance of predictors that utilize different combinations of features and how this performance is affected by structural and sequence properties of proteins. Our results indicate that an SVM predictor based on a properly scaled profile of evolutionary conservation in the form of a position specific scoring matrix (PSSM) significantly outperforms a PSSM-based neural network predictor. The highest accuracy is achieved by SVM predictor that combines the profile of evolutionary conservation with low-resolution structural information. Our results also show that knowledge-based predictors of DNA-binding sites perform significantly better on proteins from mainly-alpha structural class and that the performance of these predictors is significantly correlated with certain structural and sequence properties of proteins. These observations suggest that it may be possible to assign a reliability index to the overall accuracy of the prediction of DNA-binding sites in any given protein using its sequence and structural properties. A web-server implementation of the predictors is freely available online at http://lcg.rit.albany.edu/dp-bind/.

Amino Acid Sequence↗

The infant as a prelinguistic model for language learning impairments: predicting from event-related potentials to behavior.

Associations between efficient processing of brief, rapidly presented, successive stimuli and language learning impairments (LLI) in older children and adults have been well documented. In this paper we examine the role that impaired rapid auditory processing (RAP) might play during early language acquisition. Using behavioral measures we have demonstrated that RAP abilities in infancy are critically linked to later language abilities for both non-speech and speech stimuli. Variance in infant RAP thresholds reliably predict language outcome at 3 years-of-age for infants at risk for LLI and control infants. We present data here describing patterns of electrocortical (EEG/ERP) activation at 6 month-of-age to the same non-verbal stimuli used in our behavioral studies. Well-defined differences were seen between infants from families with a history of LLI (FH+) and FH- controls in the amplitude of the mismatch response (MMR) as well as the latency of the N250 component in the 70 ms ISI condition only. Smaller mismatch responses and delayed onsets of the N250 component were seen in the FH+ group. The latency differences in the N250 component, but not the MMR amplitude variation, were significantly related to 24-month language outcome. Such converging tasks provide the opportunity to examine early precursors of LLI and allow the opportunity for earlier identification and intervention.

Brain Mapping↗

A canonical correlation neural network for multicollinearity and functional data.

We review a recent neural implementation of Canonical Correlation Analysis and show, using ideas suggested by Ridge Regression, how to make the algorithm robust. The network is shown to operate on data sets which exhibit multicollinearity. We develop a second model which not only performs as well on multicollinear data but also on general data sets. This model allows us to vary a single parameter so that the network is capable of performing Partial Least Squares regression (at one extreme) to Canonical Correlation Analysis (at the other)and every intermediate operation between the two. On multicollinear data, the parameter setting is shown to be important but on more general data no particular parameter setting is required. Finally, we develop a second penalty term which acts on such data as a smoother in that the resulting weight vectors are much smoother and more interpretable than the weights without the robustification term. We illustrate our algorithms on both artificial and real data.

Child, Preschool↗