PubMed Health⌕ Search

Biomedical subjects

Hsien-Da Huang

Publications and source records attributed to Hsien-Da Huang.

At least 19 recordsLinked to original sources

RINGdb: an integrated database for G protein-coupled receptors and regulators of G protein signaling.

BACKGROUND: Many marketed therapeutic agents have been developed to modulate the function of G protein-coupled receptors (GPCRs). The regulators of G-protein signaling (RGS proteins) are also being examined as potential drug targets. To facilitate clinical and pharmacological research, we have developed a novel integrated biological database called RINGdb to provide comprehensive and organized RGS protein and GPCR information. RESULTS: RINGdb contains information on mutations, tissue distributions, protein-protein interactions, diseases/disorders and other features, which has been automatically collected from the Internet and manually extracted from the literature. In addition, RINGdb offers various user-friendly query functions to answer different questions about RGS proteins and GPCRs such as their possible contribution to disease processes, the putative direct or indirect relationship between RGS proteins and GPCRs. RINGdb also integrates organized database cross-references to allow users direct access to detailed information. The database is now available at http://ringdb.csie.ncu.edu.tw/ringdb/. CONCLUSION: RINGdb is the only integrated database on the Internet to provide comprehensive RGS protein and GPCR information. This knowledge base will be useful for clinical research, drug discovery and GPCR signaling pathway research.

Amino Acid Sequence↗

ViTa: prediction of host microRNAs targets on viruses.

MicroRNAs (miRNAs) are involved in various biological processes by suppressing gene expression. A recent work has indicated that host miRNAs are also capable of regulating viral gene expression by targeting the virus genomes. To investigate regulatory relationships between host miRNAs and related viruses, we present a novel database, namely ViTa, to curate the known virus miRNA genes and the known/putative target sites of human, mice, rat and chicken miRNAs. Known miRNAs are obtained from miRBase. Virus data are collected and referred from ICTVdB, VBRC and VirGen. Experimentally validated miRNA targets on viruses were derived from literatures. Then, miRanda and TargetScan are utilized to predict miRNA targets within virus genomes. ViTa also provides the virus annotations, virus-infected tissues and tissue specificity of host miRNAs. This work also facilitates the comparisons between subtypes of viruses, such as influenza viruses, human liver viruses and the conserved regions between viruses. Both textual and graphical web interfaces are provided to facilitate the data retrieves in the ViTa database. The database is now freely available at http://vita.mbc.nctu.edu.tw/.

Animals↗

An agent-based system to discover protein-protein interactions, identify protein complexes and proteins with multiple peptide mass fingerprints.

Proteins "work together" by actually binding to form multicomponent complexes that carry out specific functions. Proteomic analyses based on the mass spectrum are now key methods to determine the components in protein complexes. The protein-protein interaction or functional association may be known to exist among the extracted protein spots while analyzing the proteins on the 2D gel. In this study, we develop an agent-based system, namely AgentMultiProtIdent, which integrated two protein identification tools and a variety of databases storing relations among proteins and used to discover protein-protein interactions and protein functional associations, and identify protein complexes and proteins with multiple peptide mass fingerprints as input. The system takes Multiple Peptide Mass Fingerprints (PMFs) as a whole in the protein complex or protein identification. With the relations among proteins, it may greatly improve the accuracy of identification of protein complexes. Also, possible relationship of the multiple peptide mass fingerprints, such as ontology relation, can be discovered by our system, especially in the identification of protein complexes. The agent-based system is now available on the Web at http://dbms104.csie.ncu.edu.tw/ approximately protein/NEW2/.

Animals↗

RNAMST: efficient and flexible approach for identifying RNA structural homologs.

RNA molecules fold into characteristic secondary structures for their diverse functional activities such as post-translational regulation of gene expression. Searching homologs of a pre-defined RNA structural motif, which may be a known functional element or a putative RNA structural motif, can provide useful information for deciphering RNA regulatory mechanisms. Since searching for the RNA structural homologs among the numerous RNA sequences is extremely time-consuming, this work develops a data preprocessing strategy to enhance the search efficiency and presents RNAMST, which is an efficient and flexible web server for rapidly identifying homologs of a pre-defined RNA structural motif among numerous RNA sequences. Intuitive user interface are provided on the web server to facilitate the predictive analysis. By comparing the proposed web server to other tools developed previously, RNAMST performs remarkably more efficiently and provides more effective and flexible functions. RNAMST is now available on the web at http://bioinfo.csie.ncu.edu.tw/~rnamst/.

Internet↗

ProKware: integrated software for presenting protein structural properties in protein tertiary structures.

Protein tertiary structure plays an essential role in deciphering protein functions, especially protein structural properties, including domains, active sites and post-translational modifications. These properties typically yield useful clues for understanding protein functions. This work presents an integrated software, named ProKware, that presents protein structural properties in protein tertiary structures, such as domains, functional sites, families, active sites, binding sites, post-translational modifications and domain-domain interaction. Using this web-based and Windows-based interface, users can manipulate and visualize three-dimensional protein structures, as well as the supported structural properties that are curated in the protein knowledge database. ProKware is an effective and convenient solution for investigating protein functions and structural relationships. This software can be accessed on the internet at http://ProKware.mbc.nctu.edu.tw/.

Binding Sites↗

RegRNA: an integrated web server for identifying regulatory RNA motifs and elements.

Numerous regulatory structural motifs have been identified as playing essential roles in transcriptional and post-transcriptional regulation of gene expression. RegRNA is an integrated web server for identifying the homologs of regulatory RNA motifs and elements against an input mRNA sequence. Both sequence homologs and structural homologs of regulatory RNA motifs can be recognized. The regulatory RNA motifs supported in RegRNA are categorized into several classes: (i) motifs in mRNA 5'-untranslated region (5'-UTR) and 3'-UTR; (ii) motifs involved in mRNA splicing; (iii) motifs involved in transcriptional regulation; (iv) riboswitches; (v) splicing donor/acceptor sites; (vi) inverted repeats; and (vii) miRNA target sites. The experimentally validated regulatory RNA motifs are extracted from literature survey and several regulatory RNA motif databases, such as UTRdb, TRANSFAC, alternative splicing database (ASD) and miRBase. A variety of computational programs are integrated for identifying the homologs of the regulatory RNA motifs. An intuitive user interface is designed to facilitate the comprehensive annotation of user-submitted mRNA sequences. The RegRNA web server is now available at http://RegRNA.mbc.NCTU.edu.tw/.

Animals↗

Detection of discriminative sequence motifs in proteins obtained from prokaryotes grown at various temperatures.

Recent investigations on the stability of proteins have demonstrated various structural factors, but few have considered sequence factors such as protein motifs. These motifs represent highly conserved regions and describe critical regions that may only exist on proteins that remain functional at high temperatures. This investigation presents a method for identifying and comparing corresponding mesophilic and thermophilic sequence motifs between protein families. Discriminative motifs that are conserved only in the mesophilic or thermophilic subfamily are identified. Analysis of the results shows that, although the subfamilies of most protein families share similar motifs, some discriminative motifs are present in particular thermophilic/mesophilic subfamilies. The thermophilic discriminative motifs are conserved only in thermophilic organisms, revealing that physiochemical principles support thermostability.

Amino Acid Motifs↗

dbPTM: an information repository of protein post-translational modification.

dbPTM is a database that compiles information on protein post-translational modifications (PTMs), such as the catalytic sites, solvent accessibility of amino acid residues, protein secondary and tertiary structures, protein domains and protein variations. The database includes all of the experimentally validated PTM sites from Swiss-Prot, PhosphoELM and O-GLYCBASE. Only a small fraction of Swiss-Prot proteins are annotated with experimentally verified PTM. Although the Swiss-Prot provides rich information about the PTM, other structural properties and functional information of proteins are also essential for elucidating protein mechanisms. The dbPTM systematically identifies three major types of protein PTM (phosphorylation, glycosylation and sulfation) sites against Swiss-Prot proteins by refining our previously developed prediction tool, KinasePhos (http://kinasephos.mbc.nctu.edu.tw/). Solvent accessibility and secondary structure of residues are also computationally predicted and are mapped to the PTM sites. The resource is now freely available at http://dbPTM.mbc.nctu.edu.tw/.

Amino Acids↗

miRNAMap: genomic maps of microRNA genes and their target genes in mammalian genomes.

Recent work has demonstrated that microRNAs (miRNAs) are involved in critical biological processes by suppressing the translation of coding genes. This work develops an integrated database, miRNAMap, to store the known miRNA genes, the putative miRNA genes, the known miRNA targets and the putative miRNA targets. The known miRNA genes in four mammalian genomes such as human, mouse, rat and dog are obtained from miRBase, and experimentally validated miRNA targets are identified in a survey of the literature. Putative miRNA precursors were identified by RNAz, which is a non-coding RNA prediction tool based on comparative sequence analysis. The mature miRNA of the putative miRNA genes is accurately determined using a machine learning approach, mmiRNA. Then, miRanda was applied to predict the miRNA targets within the conserved regions in 3'-UTR of the genes in the four mammalian genomes. The miRNAMap also provides the expression profiles of the known miRNAs, cross-species comparisons, gene annotations and cross-links to other biological databases. Both textual and graphical web interface are provided to facilitate the retrieval of data from the miRNAMap. The database is freely available at http://mirnamap.mbc.nctu.edu.tw/.

Animals↗

Biological data warehousing system for identifying transcriptional regulatory sites from gene expressions of microarray data.

Identification of transcriptional regulatory sites plays an important role in the investigation of gene regulation. For this propose, we designed and implemented a data warehouse to integrate multiple heterogeneous biological data sources with data types such as text-file, XML, image, MySQL database model, and Oracle database model. The utility of the biological data warehouse in predicting transcriptional regulatory sites of coregulated genes was explored using a synexpression group derived from a microarray study. Both of the binding sites of known transcription factors and predicted over-represented (OR) oligonucleotides were demonstrated for the gene group. The potential biological roles of both known nucleotides and one OR nucleotide were demonstrated using bioassays. Therefore, the results from the wet-lab experiments reinforce the power and utility of the data warehouse as an approach to the genome-wide search for important transcription regulatory elements that are the key to many complex biological systems.

Algorithms↗

Database to dynamically aid probe design for virus identification.

Viral infection poses a major problem for public health, horticulture, and animal husbandry, possibly causing severe health crises and economic losses. Viral infections can be identified by the specific detection of viral sequences in many ways. The microarray approach not only tolerates sequence variations of newly evolved virus strains, but can also simultaneously diagnose many viral sequences. Many chips have so far been designed for clinical use. Most are designed for special purposes, such as typing enterovirus infection, and compare fewer than 30 different viral sequences. None considers primer design, increasing the likelihood of cross hybridization to similar sequences from other viruses. To prevent this possibility, this work establishes a platform and database that provides users with specific probes of all known viral genome sequences to facilitate the design of diagnostic chips. This work develops a system for designing probes online. A user can select any number of different viruses and set the experimental conditions such as melting temperature and length of probe. The system then returns the optimal sequences from the database. We have also developed a heuristic algorithm to calculate the probe correctness and show the correctness of the algorithm. (The system that supports probe design for identifying viruses has been published on our web page http://bioinfo.csie.ncu.edu.tw/.)

Algorithms↗

Incorporating hidden Markov models for identifying protein kinase-specific phosphorylation sites.

Protein phosphorylation, which is an important mechanism in posttranslational modification, affects essential cellular processes such as metabolism, cell signaling, differentiation, and membrane transportation. Proteins are phosphorylated by a variety of protein kinases. In this investigation, we develop a novel tool to computationally predict catalytic kinase-specific phosphorylation sites. The known phosphorylation sites from public domain data sources are categorized by their annotated protein kinases. Based on the concepts of profile Hidden Markov Models (HMM), computational models are trained from the kinase-specific groups of phosphorylation sites. After evaluating the trained models, we select the model with highest accuracy in each kinase-specific group and provide a Web-based prediction tool for identifying protein phosphorylation sites. The main contribution here is that we have developed a kinase-specific phosphorylation site prediction tool with both high sensitivity and specificity.

Computational Biology↗

KinasePhos: a web tool for identifying protein kinase-specific phosphorylation sites.

KinasePhos is a novel web server for computationally identifying catalytic kinase-specific phosphorylation sites. The known phosphorylation sites from public domain data sources are categorized by their annotated protein kinases. Based on the profile hidden Markov model, computational models are learned from the kinase-specific groups of the phosphorylation sites. After evaluating the learned models, the model with highest accuracy was selected from each kinase-specific group, for use in a web-based prediction tool for identifying protein phosphorylation sites. Therefore, this work developed a kinase-specific phosphorylation site prediction tool with both high sensitivity and specificity. The prediction tool is freely available at http://KinasePhos.mbc.nctu.edu.tw/.

Computational Biology↗

SpliceInfo: an information repository for mRNA alternative splicing in human genome.

We have developed an information repository named SpliceInfo to collect the occurrences of the four major alternative-splicing (AS) modes in human genome; these include exon skipping, 5'-alternative splicing, 3'-alternative splicing and intron retention. The dataset is derived by comparing the nucleotide and protein sequences available for a given gene for evidence of AS. Additional features such as the tissue specificity of the mRNA, the protein domain contained by exons, the GC-ratio of exons, the repeats contained within the exons, and the Gene Ontology are annotated computationally for each exonic region that is alternatively spliced. Motivated by a previous investigation of AS-related motifs such as exonic splicing enhancer and exonic splicing silencer, this resource also provides a means of identifying motifs candidates and this should help to identify potential regulatory mechanisms within a particular exonic sequence set and its two flanking intronic sequence sets. This is carried out using motif discovery tools to identify motif candidates related to alternative splicing regulation and together with a secondary structure prediction tool, will help in the identification of the structural properties of such regulatory motifs. The integrated resource is now available on http://SpliceInfo.mbc.NCTU.edu.tw/.

Alternative Splicing↗

i-Genome: a database to summarize oligonucleotide data in genomes.

BACKGROUND: Information on the occurrence of sequence features in genomes is crucial to comparative genomics, evolutionary analysis, the analyses of regulatory sequences and the quantitative evaluation of sequences. Computing the frequencies and the occurrences of a pattern in complete genomes is time-consuming. RESULTS: The proposed database provides information about sequence features generated by exhaustively computing the sequences of the complete genome. The repetitive elements in the eukaryotic genomes, such as LINEs, SINEs, Alu and LTR, are obtained from Repbase. The database supports various complete genomes including human, yeast, worm, and 128 microbial genomes. CONCLUSIONS: This investigation presents and implements an efficiently computational approach to accumulate the occurrences of the oligonucleotides or patterns in complete genomes. A database is established to maintain the information of the sequence features, including the distributions of oligonucleotide, the gene distribution, the distribution of repetitive elements in genomes and the occurrences of the oligonucleotides. The database can provide more effective and efficient way to access the repetitive features in genomes.

Alu Elements↗

Identifying transcriptional regulatory sites in the human genome using an integrated system.

This work develops an integrated system which, after a set of genes are inputted, is able to predict transcriptional regulatory sites and to detect the co-occurrence of these regulatory sites. The system integrates several site detection methods such as known site matching, over-presented oligonucleotide detection and DNA motif discovery programs. User profiles and history pages enable users to trace the sequence analyses of these transcriptional regulatory sites. Two groups of co-regulated genes were used to test the proposed system. The results predicted by the proposed system consist of known site homologs and putative regulatory sites. By comparing these sites with previously published results, the proposed system is able to help biologists identify possible candidates for the regulatory sites from groups of co-regulated genes. The integrated system is now available at http://rgsminer. csie.ncu.edu.tw/.

Databases, Nucleic Acid↗

Identifying the combination of genetic factors that determine susceptibility to cervical cancer.

Cervical cancer is common among women all over the world. Although infection with high-risk types of human papillomavirus (HPV) has been identified as the primary cause of cervical cancer, only some of those infected go on to develop cervical cancer. Obviously, the progression from HPV infection to cancer involves other environmental and host factors. Recent population-based twin and family studies have demonstrated the importance of the hereditary component of cervical cancer, associated with genetic susceptibility. Consequently, single-nucleotide polymorphism (SNP) markers and microsatellites should be considered genetic factors for determining what combinations of genetic factors are involved in precancerous changes to cervical cancer. This study employs a Bayesian network and four different decision tree algorithms, and compares the performance of these learning algorithms. The results of this study raise the possibility of investigations that could identify combinations of genetic factors, such as SNPs and microsatellites, that influence the risk associated with common complex multifactorial diseases, such as cervical cancer.

Algorithms↗

A probabilistic method to correlate ion pairs with protein thermostability.

Recent developments in research on the stability of proteins - specifically, comparisons of the ion pairs of homologous structures - show that ion pairs potentially contribute to the thermostability of proteins. This study proposes a probabilistic Bayesian statistical method to efficiently predict the thermostability of proteins based on the properties of ion pairs. The experimental results suggest that the numbers, types and bond lengths of ion pairs can be used to predict with high accuracy (up to 80%) the thermostability of functionally similar proteins. The predictions have high precision (99%), especially for hyperthermophilic proteins. Results for proteins with differing functions also indicate that the number of ion pairs is related to the thermostability of proteins, and that predictions of thermostability can also be made for proteins with different functions.

Algorithms↗