PubMed Health⌕ Search

Biomedical subjects

Chi-Ren Shyu

Publications and source records attributed to Chi-Ren Shyu.

9 recordsLinked to original sources

A fast SCOP fold classification system using content-based E-Predict algorithm.

BACKGROUND: Domain experts manually construct the Structural Classification of Protein (SCOP) database to categorize and compare protein structures. Even though using the SCOP database is believed to be more reliable than classification results from other methods, it is labor intensive. To mimic human classification processes, we develop an automatic SCOP fold classification system to assign possible known SCOP folds and recognize novel folds for newly-discovered proteins. RESULTS: With a sufficient amount of ground truth data, our system is able to assign the known folds for newly-discovered proteins in the latest SCOP v1.69 release with 92.17% accuracy. Our system also recognizes the novel folds with 89.27% accuracy using 10 fold cross validation. The average response time for proteins with 500 and 1409 amino acids to complete the classification process is 4.1 and 17.4 seconds, respectively. By comparison with several structural alignment algorithms, our approach outperforms previous methods on both the classification accuracy and efficiency. CONCLUSION: In this paper, we build an advanced, non-parametric classifier to accelerate the manual classification processes of SCOP. With satisfactory ground truth data from the SCOP database, our approach identifies relevant domain knowledge and yields reasonably accurate classifications. Our system is publicly accessible at http://ProteinDBS.rnet.missouri.edu/E-Predict.php.

Algorithms↗

Refined repetitive sequence searches utilizing a fast hash function and cross species information retrievals.

BACKGROUND: Searching for small tandem/disperse repetitive DNA sequences streamlines many biomedical research processes. For instance, whole genomic array analysis in yeast has revealed 22 PHO-regulated genes. The promoter regions of all but one of them contain at least one of the two core Pho4p binding sites, CACGTG and CACGTT. In humans, microsatellites play a role in a number of rare neurodegenerative diseases such as spinocerebellar ataxia type 1 (SCA1). SCA1 is a hereditary neurodegenerative disease caused by an expanded CAG repeat in the coding sequence of the gene. In bacterial pathogens, microsatellites are proposed to regulate expression of some virulence factors. For example, bacteria commonly generate intra-strain diversity through phase variation which is strongly associated with virulence determinants. A recent analysis of the complete sequences of the Helicobacter pylori strains 26695 and J99 has identified 46 putative phase-variable genes among the two genomes through their association with homopolymeric tracts and dinucleotide repeats. Life scientists are increasingly interested in studying the function of small sequences of DNA. However, current search algorithms often generate thousands of matches -- most of which are irrelevant to the researcher. RESULTS: We present our hash function as well as our search algorithm to locate small sequences of DNA within multiple genomes. Our system applies information retrieval algorithms to discover knowledge of cross-species conservation of repeat sequences. We discuss our incorporation of the Gene Ontology (GO) database into these algorithms. We conduct an exhaustive time analysis of our system for various repetitive sequence lengths. For instance, a search for eight bases of sequence within 3.224 GBases on 49 different chromosomes takes 1.147 seconds on average. To illustrate the relevance of the search results, we conduct a search with and without added annotation terms for the yeast Pho4p binding sites, CACGTG and CACGTT. Also, a cross-species search is presented to illustrate how potential hidden correlations in genomic data can be quickly discerned. The findings in one species are used as a catalyst to discover something new in another species. These experiments also demonstrate that our system performs well while searching multiple genomes -- without the main memory constraints present in other systems. CONCLUSION: We present a time-efficient algorithm to locate small segments of DNA and concurrently to search the annotation data accompanying the sequence. Genome-wide searches for short sequences often return hundreds of hits. Our experiments show that subsequently searching the annotation data can refine and focus the results for the user. Our algorithms are also space-efficient in terms of main memory requirements. Source code is available upon request.

Algorithms↗

Knowledge representation and sharing using visual semantic modeling for diagnostic medical image databases.

Information technology offers great opportunities for supporting radiologists' expertise in decision support and training. However, this task is challenging due to difficulties in articulating and modeling visual patterns of abnormalities in a computational way. To address these issues, well established approaches to content management and image retrieval have been studied and applied to assist physicians in diagnoses. Unfortunately, most of the studies lack the flexibility of sharing both explicit and tacit knowledge involved in the decision making process, while adapting to each individual's opinion. In this paper, we propose a knowledge repository and exchange framework for diagnostic image databases called "evolutionary system for semantic exchange of information in collaborative environments" (Essence). This framework uses semantic methods to describe visual abnormalities, and offers a solution for tacit knowledge elicitation and exchange in the medical domain. Also, our approach provides a computational and visual mechanism for associating synonymous semantics of visual abnormalities. We conducted several experiments to demonstrate the system's capability of matching synonym terms, and the benefit of using tacit knowledge in improving the meaningfulness of semantic queries.

Artificial Intelligence↗

ProteinDBS: a real-time retrieval system for protein structure comparison.

We have developed a web server (ProteinDBS) for the life science community to search for similar protein tertiary structures in real time. This system applies computer visualization techniques to extract the predominant visual patterns encoded in two-dimensional distance matrices generated from the three-dimensional coordinates of protein chains. When meaningful contents, represented in a multi-dimensional feature space, have been extracted from distance matrices, an advanced indexing structure, Entropy Balanced Statistical (EBS) k-d tree, is utilized to index the data. Our system is able to return search results in ranked order from a database with 46 075 chains in seconds, exhibiting a reasonably high degree of precision. To our knowledge, this is the first real-time search engine for protein structure comparison. ProteinDBS provides two types of query method: query by Protein Data Bank protein chain ID and by new structures uploaded by users. The system is hosted at http://ProteinDBS.rnet.missouri.edu.

Computer Graphics↗

ACMES: fast multiple-genome searches for short repeat sequences with concurrent cross-species information retrieval.

We have developed a web server for the life sciences community to use to search for short repeats of DNA sequence of length between 3 and 10,000 bases within multiple species. This search employs a unique and fast hash function approach. Our system also applies information retrieval algorithms to discover knowledge of cross-species conservation of repeat sequences. Furthermore, we have incorporated a part of the Gene Ontology database into our information retrieval algorithms to broaden the coverage of the search. Our web server and tutorial can be found at http://acmes.rnet.missouri.edu.

Algorithms↗

Design and evaluation of a personal digital assistant- based alerting service for clinicians.

PURPOSE: This study describes the system architecture and user acceptance of a suite of programs that deliver information about newly updated library resources to clinicians' personal digital assistants (PDAs). DESCRIPTION: Participants received headlines delivered to their PDAs alerting them to new books, National Guideline Clearinghouse guidelines, Cochrane Reviews, and National Institutes of Health (NIH) Clinical Alerts, as well as updated content in UpToDate, Harrison's Online, Scientific American Medicine, and Clinical Evidence. Participants could request additional information for any of the headlines, and the information was delivered via e-mail during their next synchronization. Participants completed a survey at the conclusion of the study to gauge their opinions about the service. RESULTS/OUTCOME: Of the 816 headlines delivered to the 16 study participants' PDAs during the project, Scientific American Medicine generated the highest proportion of headline requests at 35%. Most users of the PDA Alerts software reported that they learned about new medical developments sooner than they otherwise would have, and half reported that they learned about developments that they would not have heard about at all. While some users liked the PDA platform for receiving headlines, it seemed that a Web database that allowed tailored searches and alerts could be configured to satisfy both PDA-oriented and e-mail-oriented users.

Attitude to Computers↗

Whither biological database research?

We consider how the landscape of biological databases may evolve in the future, and what research is needed to realize this evolution. We suggest today's dispersal of diverse resources will only increase as the number and size of those resources, driving the need for semantic interoperability even more strongly. Because the complexity of the questions biologists want answered automatically continues to rapidly escalate, we will need to draw upon high-performance computing resources such as the GRID to process complex queries. Finally, we still need data, and our ways of acquiring and curating data must improve by orders of magnitude.

Computational Biology↗

Automated storage and retrieval of thin-section CT images to assist diagnosis: system description and preliminary assessment.

A software system and database for computer-aided diagnosis with thin-section computed tomographic (CT) images of the chest was designed and implemented. When presented with an unknown query image, the system uses pattern recognition to retrieve visually similar images with known diagnoses from the database. A preliminary validation trial was conducted with 11 volunteers who were asked to select the best diagnosis for a series of test images, with and without software assistance. The percentage of correct answers increased from 29% to 62% with computer assistance. This finding suggests that this system may be useful for computer-assisted diagnosis.

Databases, Factual↗

Methylation microarray analysis of late-stage ovarian carcinomas distinguishes progression-free survival in patients and identifies candidate epigenetic markers.

PURPOSE: The purpose of this study was to profile methylation alterations of CpG islands in ovarian tumors and to identify candidate markers for diagnosis and prognosis of the disease. EXPERIMENTAL DESIGN: A global analysis of DNA methylation using a novel microarray approach called differential methylation hybridization was performed on 19 patients with stage III and IV ovarian carcinomas. RESULTS: Hierarchical clustering identified two groups of patients with distinct methylation profiles. Tumors from group 1 contained high levels of concurrent methylation, whereas group 2 tumors had lower tumor methylation levels. The duration of progression-free survival after chemotherapy was significantly shorter for patients in group 1 compared with group 2 (P < 0.001). Differential methylation in tumors was independently confirmed by methylation-specific PCR. CONCLUSIONS: The data suggest that a higher degree of CpG island methylation is associated with early disease recurrence after chemotherapy. The differential methylation hybridization assay also identified a select group of CpG island loci that are potentially useful as epigenetic markers for predicting treatment outcome in ovarian cancer patients.

Biomarkers, Tumor↗