PubMed Health⌕ Search

Biomedical subjects

Jaehyun Sim

Publications and source records attributed to Jaehyun Sim.

4 recordsLinked to original sources

The tissue microarray object model: a data model for storage, analysis, and exchange of tissue microarray experimental data.

CONTEXT: Tissue microarray (TMA) is an array-based technology allowing the examination of hundreds of tissue samples on a single slide. To handle, exchange, and disseminate TMA data, we need standard representations of the methods used, of the data generated, and of the clinical and histopathologic information related to TMA data analysis. OBJECTIVE: To create a comprehensive data model with flexibility that supports diverse experimental designs and with expressivity and extensibility that enables an adequate and comprehensive description of new clinical and histopathologic data elements. DESIGN: We designed a tissue microarray object model (TMA-OM). Both the array information and the experimental procedure models are created by referring to the microarray gene expression object model, minimum information specification for in situ hybridization and immunohistochemistry experiments, and the TMA data exchange specifications. The clinical and histopathologic information model is created by using College of American Pathologists cancer protocols and National Cancer Institute common data elements. Microarray Gene Expression Data Ontology, the Unified Medical Language System, and the terms extracted from College of American Pathologists cancer protocols and NCI common data elements are used to create a controlled vocabulary for unambiguous annotation. RESULT: The TMA-OM consists of 111 classes in 17 packages to represent clinical and histopathologic information as well as experimental data for any type of cancer. We implemented a Web-based application for TMA-OM, supporting data export in XML format conforming to the TMA data exchange specifications or the document type definition derived from TMA-OM. CONCLUSIONS: The TMA-OM provides a comprehensive data model for storage, analysis, and exchange of TMA data and facilitates model-level integration of other biological models.

Gene Expression Profiling↗

Study of protein-protein interaction using conformational space annealing.

We apply conformational space annealing (CSA), an efficient global optimization method, to the study of protein-protein interaction. The CSA is incorporated into the Tinker molecular modeling package along with a B-spline method for CAPRI Round 5 experiments. We have used an energy function for the protein-protein interaction that consists of electrostatic interaction, van der Waals interaction, and solvation energy terms represented by the occupancy desolvation method. The parameters of the AMBER94 all-atom empirical force field are used. Each energy term is calculated by precalculated grid potentials and B-spline method approximation. The ligand protein is placed inside a sphere of 50 A radius centered at an appropriate location, and the CSA rigid docking studies are carried out to find stable complexes. Up to 10 complexes are selected using the K-mean clustering method and biological information when available. These complexes are energy-minimized for further refinement by considering the flexibility of interacting proteins. The results show that the CSA method has a potential for the study of protein-protein interaction.

Algorithms↗

PPRODO: prediction of protein domain boundaries using neural networks.

Successful prediction of protein domain boundaries provides valuable information not only for the computational structure prediction of multidomain proteins but also for the experimental structure determination. Since protein sequences of multiple domains may contain much information regarding evolutionary processes such as gene-exon shuffling, this information can be detected by analyzing the position-specific scoring matrix (PSSM) generated by PSI-BLAST. We have presented a method, PPRODO (Prediction of PROtein DOmain boundaries) that predicts domain boundaries of proteins from sequence information by a neural network. The network is trained and tested using the values obtained from the PSSM generated by PSI-BLAST. A 10-fold cross-validation technique is performed to obtain the parameters of neural networks using a nonredundant set of 522 proteins containing 2 contiguous domains. PPRODO provides good and consistent results for the prediction of domain boundaries, with accuracy of about 66% using the +/-20 residue criterion. The PPRODO source code, as well as all data sets used in this work, are available from http://gene.kias.re.kr/ approximately jlee/pprodo/.

Amino Acids↗

Prediction of protein solvent accessibility using fuzzy k-nearest neighbor method.

MOTIVATION: The solvent accessibility of amino acid residues plays an important role in tertiary structure prediction, especially in the absence of significant sequence similarity of a query protein to those with known structures. The prediction of solvent accessibility is less accurate than secondary structure prediction in spite of improvements in recent researches. The k-nearest neighbor method, a simple but powerful classification algorithm, has never been applied to the prediction of solvent accessibility, although it has been used frequently for the classification of biological and medical data. RESULTS: We applied the fuzzy k-nearest neighbor method to the solvent accessibility prediction, using PSI-BLAST profiles as feature vectors, and achieved high prediction accuracies. With leave-one-out cross-validation on the ASTRAL SCOP reference dataset constructed by sequence clustering, our method achieved 64.1% accuracy for a 3-state (buried/intermediate/exposed) prediction (thresholds of 9% for buried/intermediate and 36% for intermediate/exposed) and 86.7, 82.0, 79.0 and 78.5% accuracies for 2-state (buried/exposed) predictions (thresholds of each 0, 5, 16 and 25% for buried/exposed), respectively. Our method also showed slightly better accuracies than other methods by about 2-5% on the RS126 dataset and a benchmarking dataset with 229 proteins. AVAILABILITY: Program and datasets are available at http://biocom1.ssu.ac.kr/FKNNacc/ CONTACT: jul@ssu.ac.kr.

Algorithms↗