PubMed Health⌕ Search

Biomedical subjects

Jagath C Rajapakse

Publications and source records attributed to Jagath C Rajapakse.

14 recordsLinked to original sources

Two-stage support vector regression approach for predicting accessible surface areas of amino acids.

We address the problem of predicting solvent accessible surface area (ASA) of amino acid residues in protein sequences, without classifying them into buried and exposed types. A two-stage support vector regression (SVR) approach is proposed to predict real values of ASA from the position-specific scoring matrices generated from PSI-BLAST profiles. By adding SVR as the second stage to capture the influences on the ASA value of a residue by those of its neighbors, the two-stage SVR approach achieves improvements of mean absolute errors up to 3.3%, and correlation coefficients of 0.66, 0.68, and 0.67 on the Manesh dataset of 215 proteins, the Barton dataset of 502 nonhomologous proteins, and the Carugo dataset of 338 proteins, respectively, which are better than the scores published earlier on these datasets. A Web server for protein ASA prediction by using a two-stage SVR method has been developed and is available (http://birc.ntu.edu.sg/~ pas0186457/asa.html).

Amino Acids↗

Learning functional structure from fMR images.

We propose a novel method using Bayesian networks to learn the structure of effective connectivity among brain regions involved in a functional MR experiment. The approach is exploratory in the sense that it does not require an a priori model as in the earlier approaches, such as the Structural Equation Modeling or Dynamic Causal Modeling, which can only affirm or refute the connectivity of a previously known anatomical model or a hypothesized model. The conditional probabilities that render the interactions among brain regions in Bayesian networks represent the connectivity in the complete statistical sense. The present method is applicable even when the number of regions involved in the cognitive network is large or unknown. We demonstrate the present approach by using synthetic data and fMRI data collected in silent word reading and counting Stroop tasks.

Adult↗

Contextual modeling of functional MR images with conditional random fields.

This paper presents a conditional random field (CRF) approach to fuse contextual dependencies in functional magnetic resonance imaging (fMRI) data for the detection of brain activation. The interactions among both activation (activated/inactive) labels and observed data of brain voxels are unified in a probabilistic framework based on the CRF, where the interaction strength can be adaptively adjusted in terms of the data similarity of neighboring sites. Compared to earlier detection methods, including statistical parametric mapping and Markov random field, the proposed method avoids the suppression of high frequency information and relaxes the strong assumption of conditional independence of observed data. Experimental results show that the proposed approach effectively integrates contextual constraints within the detection process and robustly detects brain activities from fMRI data.

Algorithms↗

Phonological processing in Chinese-English bilingual biscriptals: an fMRI study.

Different activation loci have been reported for language processing in unilingual Chinese and unilingual English participants, as well as in bilingual readers of English and French, two alphabetic languages. Nevertheless, the extant imaging work on Mandarin-English bilinguals favors common neural substrates for English and Chinese, languages with contrasting oral and written forms. We investigated the phonological processes in reading for English-Chinese biscriptals using a homophone matching task with parallel behavioral (n = 28) and fMRI (n = 6) experiments. Unlike previous reports, we observed distinct regions of activation for Mandarin in the left and right frontal lobes, the left temporal lobe, and the right occipital lobe, plus distinct regions of activation for English bilaterally in both the frontal and parietal lobes. The implications of these novel findings are discussed with reference to language representation in bilinguals.

Adolescent↗

Segmentation of subcortical brain structures using fuzzy templates.

We propose a novel method to automatically segment subcortical structures of human brain in magnetic resonance images by using fuzzy templates. A set of fuzzy templates of the structures based on features such as intensity, spatial location, and relative spatial relationship among structures are first created from a set of training images by defining the fuzzy membership functions and by fusing the information of features. Segmentation is performed by registering the fuzzy templates of the structures on the test image and then by fusing them with the tissue maps of the test image. The final decision is taken in order to optimize the certainty in the intensity, location, relative position, and tissue content of the structure. Our method does not require specific expert definition of each structure or manual interactions during segmentation process. The technique is demonstrated with the segmentation of five structures: thalamus, putamen, caudate, hippocampus, and amygdala; the performance of the present method is comparable with previous techniques.

Algorithms↗

Prediction of protein relative solvent accessibility with a two-stage SVM approach.

Information on relative solvent accessibility (RSA) of amino acid residues in proteins provides valuable clues to the prediction of protein structure and function. A two-stage approach with support vector machines (SVMs) is proposed, where an SVM predictor is introduced to the output of the single-stage SVM approach to take into account the contextual relationships among solvent accessibilities for the prediction. By using the position-specific scoring matrices (PSSMs) generated by PSI-BLAST, the two-stage SVM approach achieves accuracies up to 90.4% and 90.2% on the Manesh data set of 215 protein structures and the RS126 data set of 126 nonhomologous globular proteins, respectively, which are better than the highest published scores on both data sets to date. A Web server for protein RSA prediction using a two-stage SVM method has been developed and is available (http://birc.ntu.edu.sg/~pas0186457/rsa.html).

Amino Acids↗

Multiple SVM-RFE for gene selection in cancer classification with expression data.

This paper proposes a new feature selection method that uses a backward elimination procedure similar to that implemented in support vector machine recursive feature elimination (SVM-RFE). Unlike the SVM-RFE method, at each step, the proposed approach computes the feature ranking score from a statistical analysis of weight vectors of multiple linear SVMs trained on subsamples of the original training data. We tested the proposed method on four gene expression datasets for cancer classification. The results show that the proposed feature selection method selects better gene subsets than the original SVM-RFE and improves the classification accuracy. A Gene Ontology-based similarity assessment indicates that the selected subsets are functionally diverse, further validating our gene selection method. This investigation also suggests that, for gene expression-based cancer classification, average test error from multiple partitions of training and test sets can be recommended as a reference of performance quality.

Algorithms↗

Approach and applications of constrained ICA.

This paper presents the technique of constrained independent component analysis (cICA) and demonstrates two applications, less-complete ICA, and ICA with reference (ICA-R). The cICA is proposed as a general framework to incorporate additional requirements and prior information in the form of constraints into the ICA contrast function. The adaptive solutions using the Newton-like learning are proposed to solve the constrained optimization problem. The applications illustrate the versatility of the cICA by separating subspaces of independent components according to density types and extracting a set of desired sources when rough templates are available. The experiments using face images and functional MR images demonstrate the usage and efficacy of the cICA.

Algorithms↗

Proteomic cancer classification with mass spectrometry data.

The ultimate goal of cancer proteomics is to adapt proteomic technologies for routine use in clinical laboratories for the purpose of diagnostic and prognostic classification of disease states, as well as in evaluating drug toxicity and efficacy. Analysis of tumor-specific proteomic profiles may also allow better understanding of tumor development and the identification of novel targets for cancer therapy. The biological variability among patient samples as well as the huge dynamic range of biomarker concentrations are currently the main challenges facing efforts to deduce diagnostic patterns that are unique to specific disease states. While several strategies exist to address this problem, we focus here on cancer classification using mass spectrometry (MS) for proteomic profiling and biomarker identification. Recent advances in MS technology are starting to enable high-throughput profiling of the protein content of complex samples. For cancer classification, the protein samples from cancer patients and noncancer patients or from different cancer stages are analyzed through MS instruments and the MS patterns are used to build a diagnostic classifier. To illustrate the importance of feature selection in cancer classification, we present a method based on support vector machine-recursive feature elimination (SVM-RFE), demonstrated on two cancer datasets from ovarian and lung cancer.

Biomarkers, Tumor↗

Graphical approach to weak motif recognition.

We address the weak motif recognition problem in DNA sequences, which extends the general motif recognition to more difficult cases, allowing more degenerations in motif instances. Several algorithms have earlier attempted to find weak motifs in DNA sequences but with limitations. In this paper, we propose a graph-based algorithm for weak motif detection, which uses dynamic programming approach to find cliques indicating motif instances. The experiments on synthetic datasets show that the algorithm finds weak motif instances more accurately and efficiently compared to earlier approaches. Its performances on real datasets in finding transcription factor binding sites are comparable with the existing techniques.

Algorithms↗

Splice site detection with a higher-order markov model implemented on a neural network.

The performance of the ab inito gene prediction approaches mostly depends on the effectiveness of detecting the splice sites. This paper addresses the problem of splice site detection using higher-order Markov models. The tenet of our approach is to brace the higher-order dependencies of a Markov model by a neural network that receives the inputs from low-order Markov chains. The method is able not only to capture the higher-order dependencies in the bases of the consensus sequence immediately surrounding the splice site but also to distinguish the characteristics of the coding and non-coding regions on both sides of the splice site. Our experiments indicate that the present method achieves better accuracies over the techniques employing low-order Markov chains and other earlier approaches.

Base Sequence↗

Multi-class support vector machines for protein secondary structure prediction.

The solution of binary classification problems using the Support Vector Machine (SVM) method has been well developed. Though multi-class classification is typically solved by combining several binary classifiers, recently, several multi-class methods that consider all classes at once have been proposed. However, these methods require resolving a much larger optimization problem and are applicable to small datasets. Three methods based on binary classifications: one-against-all (OAA), one-against-one (OAO), and directed acyclic graph (DAG), and two approaches for multi-class problem by solving one single optimization problem, are implemented to predict protein secondary structure. Our experiments indicate that multi-class SVM methods are more suitable for protein secondary structure (PSS) prediction than the other methods, including binary SVMs, because their capacity to solve an optimization problem in one step. Furthermore, in this paper, we argue that it is feasible to extend the prediction accuracy by adding a second-stage multi-class SVM to capture the contextual information among secondary structural elements and thereby further improving the accuracies. We demonstrate that two-stage SVMs perform better than single-stage SVM techniques for PSS prediction using two datasets and report a maximum accuracy of 79.5%.

Computational Biology↗

Markov encoding for detecting signals in genomic sequences.

We present a technique to encode the inputs to neural networks for the detection of signals in genomic sequences. The encoding is based on lower-order Markov models which incorporate known biological characteristics in genomic sequences. The neural networks then learn intrinsic higher-order dependencies of nucleotides at the signal sites. We demonstrate the efficacy of the Markov encoding method in the detection of three genomic signals, namely, splice sites, transcription start sites, and translation initiation sites.

Base Sequence↗