PubMed Health⌕ Search

Biomedical subjects

Yanda Li

Publications and source records attributed to Yanda Li.

18 recordsLinked to original sources

dbRES: a web-oriented database for annotated RNA editing sites.

Although a large amount of experimentally derived information about RNA editing sites currently exists, this information has remained scattered in a variety of sources and in diverse data formats. Availability of standard collections for high-quality experimental data will be by of great help for systematic studying of RNA editing, especially for developing computational algorithm to predict RNA editing site. dbRES (http://bioinfo.au.tsinghua.edu.cn/dbRES) is a public database of known RNA editing sites. All sites are manually curated from literature and GenBank annotations. dbRES version 1.1 contains 5437 RNA editing sites of 251 transcripts, covering 96 organisms across plant, metazoan, protozoa, fungi and virus. dbRES provides comprehensive annotations and data summaries, including (but not limited to) transcript sequences, RNA editing types, editing site locations, amino acid changes, organisms, subcellular organelles (if available), cited references, etc. A user-friendly web interface is developed to facilitate both retrieving data and online display of RNA edit site information.

Databases, Nucleic Acid↗

Primary transcripts and expressions of mammal intergenic microRNAs detected by mapping ESTs to their flanking sequences.

MicroRNAs (miRNAs) are a class of approximately 22-nt small RNAs that regulate posttranscriptional gene expression. Thousands of expressed sequence tags (ESTs) have been identified by using upstream 2500-nt and downstream 4000-nt flanking sequences to BLAST in the dbEST database. The cotranscription of the miRNAs and their flanking sequences covered by the matched ESTs is verified by RT-PCR. It directly reveals that a large portion of mammalian intergenic miRNAs are first transcribed as long primary transcripts (pri-miRNAs). Also, the transcripts' ranges of tens of pri-miRNAs are predicted by the EST-extension method. We then extracted the tissue-specific expression information from the annotations of the matched ESTs and established the expression profile of the studied miRNAs for tens of tissues. This provided a new way to establish the expression profiles of miRNAs. Results show that the human brain, lung, liver, and eye and the mouse brain, eye, and mammary gland are tissues in which enriched numbers of miRNAs are expressed.

3' Flanking Region↗

Alteration of protein subcellular location and domain formation by alternative translational initiation.

Alternative translation is an important cellular mechanism contributing to the generation of proteins and the diversity of protein functions. Instead of studying individual cases, we systematically analyzed the alteration of protein subcellular location and domain formation by alternative translational initiation in eukaryotes. The results revealed that 85.7% of alternative translation events generated biological diversity, attributed to different subcellular localizations and distinct domain contents in alternative isoforms. Analysis of isoelectric point values revealed that most N-terminal truncated isoforms significantly lowered their isoelectric point values targeted at different subcellular localizations, whereas they had conserved domain contents the same as the full-length isoforms. Furthermore, Fisher's exact test indicated that the two ways-targeting at different cellular compartments and changing domain contents-were negatively associated. The N-term truncated isoforms should have only one way to diversify their functions distinct from the full-length ones. The peculiar consequence of subcellular relocation as well as change of domain contents reflected the very high level of biological complexity as alternative usage of initiation codons.

Animals↗

Automatic removal of the eye blink artifact from EEG using an ICA-based template matching approach.

Independent component analysis (ICA) proves to be effective in the removing the ocular artifact from electroencephalogram recordings (EEG). While using ICA in ocular artifact correction, a crucial step is to correctly identify the artifact components among the decomposed independent components. In most previous works, this step of selecting the artifact components was manually implemented, which is time consuming and inconvenient when dealing with a large amount of EEG data. We present a new method which automatically selects the eye blink artifact components based on the pattern of their scalp topographies, which can be exemplified as a template matching approach. The feasibility of using a fixed template for singling out the eye blink component after ICA decomposition was validated by an experiment in which 18 subjects among the 21 subjects involved exhibited a highly consistent pattern of eye blink scalp topographies. Since only the spatial feature is employed for singling out the eye blink component, the proposed method is very efficient and easy to implement. Objective evaluation of the real results shows that the proposed algorithm can remove the eye blink artifact from the EEG while causing little distortion to the underlying brain activities.

Algorithms↗

Stomatin-like protein 2 is overexpressed in cancer and involved in regulating cell growth and cell adhesion in human esophageal squamous cell carcinoma.

PURPOSE: Stomatin-like protein 2 (SLP-2) is a novel and unusual stomatin homologue of unknown functions. It has been implicated in interaction with erythrocyte cytoskeleton and presumably other integral membrane proteins, but not directly with the membrane bilayer. We show here the involvement of SLP-2 in human esophageal squamous cell carcinoma (ESCC), lung cancer, laryngeal cancer, and endometrial adenocarcinoma and the effects of SLP-2 on ESCC cells. EXPERIMENTAL DESIGN: Previous work of cDNA microarray in our laboratory revealed that SLP-2 was significantly up-regulated in ESCC. The expression of SLP-2 was further evaluated in human ESCC, lung cancer, laryngeal cancer, and endometrial adenocarcinoma by semiquantitative reverse transcription-PCR, Western blot, and immunohistochemistry. Mutation detection of SLP-2 exons was done by PCR and automated sequencing. Antisense SLP-2 eukaryotic expression plasmids were constructed and transfected into human ESCC cell line KYSE450. 3-(4,5-Dimethylthiazol-2-yl)-2,5-diphenyltetrazolium bromide assay, clonogenecity assay, flow cytometry assay, nude mice tumorigenetic assay, and cell attachment assay were done to investigate the roles of SLP-2 gene. RESULTS: All tumor types we tested showed overexpression of SLP-2 compared with their normal counterparts (P < or = 0.05). Moreover, immunohistochemistry analysis of mild dysplasia, severe dysplasia, and ESCC showed that overexpression of SLP-2 occurred in premalignant lesions. Mutation analysis indicated that no mutation was found in SLP-2 exons. KYSE450 cells transfected with antisense SLP-2 showed decreased cell growth, proliferation, tumorigenecity, and cell adhesion. CONCLUSIONS: SLP-2 was first identified as a novel cancer-related gene overexpressed in human ESCC, lung cancer, laryngeal cancer, and endometrial adenocarcinoma. Decreased cell growth, cell adhesion, and tumorigenesis in the antisense transfectants revealed that SLP-2 may be important in tumorigenesis.

Adenocarcinoma↗

Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine.

BACKGROUND: MicroRNAs (miRNAs) are a group of short (approximately 22 nt) non-coding RNAs that play important regulatory roles. MiRNA precursors (pre-miRNAs) are characterized by their hairpin structures. However, a large amount of similar hairpins can be folded in many genomes. Almost all current methods for computational prediction of miRNAs use comparative genomic approaches to identify putative pre-miRNAs from candidate hairpins. Ab initio method for distinguishing pre-miRNAs from sequence segments with pre-miRNA-like hairpin structures is lacking. Being able to classify real vs. pseudo pre-miRNAs is important both for understanding of the nature of miRNAs and for developing ab initio prediction methods that can discovery new miRNAs without known homology. RESULTS: A set of novel features of local contiguous structure-sequence information is proposed for distinguishing the hairpins of real pre-miRNAs and pseudo pre-miRNAs. Support vector machine (SVM) is applied on these features to classify real vs. pseudo pre-miRNAs, achieving about 90% accuracy on human data. Remarkably, the SVM classifier built on human data can correctly identify up to 90% of the pre-miRNAs from other species, including plants and virus, without utilizing any comparative genomics information. CONCLUSION: The local structure-sequence features reflect discriminative and conserved characteristics of miRNAs, and the successful ab initio classification of real and pseudo pre-miRNAs opens a new approach for discovering new miRNAs.

Animals↗

Multi-locus penetrance variance analysis method for association study in complex diseases.

Common heritable diseases often result from the action of several different genes, each of which contributes to the total observed variability in the disease trait. Traditional single-locus association approaches rely heavily on the marginal effects of single-locus and tend to ignore the multigenic nature of complex diseases. The increasing request for localizing genes underlying traits in multi-gene diseases has led to the development of some statistical methods. In this study, we develop a multi-locus analysis method - multi-locus penetrance variance analysis (MPVA), and conduct systematical simulation studies to evaluate its performance. Our results show that compared with other multi-locus methods, MPVA has some advantage in detecting complicated interactions under different epistatic models, and its performance is stable and robust.

Analysis of Variance↗

The relationship among gene expression, folding free energy and codon usage bias in Escherichia coli.

Taking advantage of microarray data in Escherichia coli genome, the relationship among mRNA expression levels, folding free energy and codon usage bias are investigated. Our results indicate that mRNA expression is correlated to the stability of mRNA secondary structure and the codon usage bias. The decrease of the stability of mRNA structure contributes to the increase of mRNA expression. There is a negative correlation between codon adaptation index (CAI) and mRNA expression in genes with less stable structure. The relationship between the stability of mRNA structure and mRNA half-life indicates the stability of mRNA structure is different from mRNA half-life.

Codon↗

ATID: a web-oriented database for collection of publicly available alternative translational initiation events.

SUMMARY: Alternative translational initiation is an important cellular mechanism contributing to the diversity of protein products and functions. We develop a database that provides a comprehensive collection of alternative translational initiation events. The purpose of this alternative translational initiation database (ATID) is to facilitate the systematic study of alternative translational initiation of genes. The current version of database contains 300 genes from Homo sapiens, Mus musculus and other species. Each of the genes has two or more isoforms due to alternative translational initiation. Resources in ATID, including gene information, alternative products of genes and domain structures of isoforms, are provided through a user-friendly web interface. AVAILABILITY: The ATID database is available for public use at http://bioinfo.au.tsinghua.edu.cn/atie/.

Amino Acid Sequence↗

The effect of U1 snRNA binding free energy on the selection of 5' splice sites.

The importance of U1 snRNA binding free energy in the regulation of alternative splicing has been studied in some genes with site-directed mutagenesis. Here we report a large-scale analysis of its impact on 5' splice site (5'ss) selection in human genome. The results show that free energy exerts different effects on alternative 5'ss choice in different situations and -8.1 kcal/mol is a threshold. When both free energies of two competing 5'ss are larger than -8.1 kcal/mol, the 5'ss with lower free energy is more frequently used. However, in other pairs of 5'ss, lower-free-energy 5'ss does not seem to be favored and even the other 5'ss is used more frequently, which suggests that very low binding free energy would impair splicing. Some observations hold true only for those alternative 5' splicing with short alternative exons (<50nt), which implies a complex mechanism of 5'ss selection involving both U1 snRNA binding free energy and regulatory factors.

Alternative Splicing↗

MicroRNA identification based on sequence and structure alignment.

MOTIVATION: MicroRNAs (miRNA) are approximately 22 nt long non-coding RNAs that are derived from larger hairpin RNA precursors and play important regulatory roles in both animals and plants. The short length of the miRNA sequences and relatively low conservation of pre-miRNA sequences restrict the conventional sequence-alignment-based methods to finding only relatively close homologs. On the other hand, it has been reported that miRNA genes are more conserved in the secondary structure rather than in primary sequences. Therefore, secondary structural features should be more fully exploited in the homologue search for new miRNA genes. RESULTS: In this paper, we present a novel genome-wide computational approach to detect miRNAs in animals based on both sequence and structure alignment. Experiments show this approach has higher sensitivity and comparable specificity than other reported homologue searching methods. We applied this method on Anopheles gambiae and detected 59 new miRNA genes. AVAILABILITY: This program is available at http://bioinfo.au.tsinghua.edu.cn/miralign. SUPPLEMENTARY INFORMATION: Supplementary information is available at http://bioinfo.au.tsinghua.edu.cn/miralign/supplementary.htm.

Algorithms↗

Multi-locus association study of schizophrenia susceptibility genes with a posterior probability method.

Schizophrenia is a serious neuropsychiatric illness affecting about 1% of the world's population. It is considered a complex inheritance disorder. A number of genes are involved in combination in the etiology of the disorder. Evidence implicates the altered dopaminergic transmission in schizophrenia. In the present study, in order to identify susceptibility genes for schizophrenia in dopaminergic metabolism, we analyzed 59 single nucleotide polymorphisms (SNPs) in 24 genes of the dopaminergic pathway among 82 unrelated patients with schizophrenia and 108 matched normal controls. Considering that traditional single-locus association studies ignore the multigenic nature of complex diseases and do not take into account possible interactions between susceptibility genes, we proposed a multi-locus analysis method, using the posterior probability of morbidity as a measure of absolute disease risk for a multi-locus genotype combination, and developed an algorithm based on perturbation and average to detect the susceptibility multi-locus genotype combinations, as well as to repress noise and avoid false positive results at our best. A three-locus SNP genotype combination involved in the interactions of COMT and ALDH3B1 genes was detected to be significantly susceptible to schizophrenia.

Aldehyde Dehydrogenase↗

HMMGEP: clustering gene expression data using hidden Markov models.

SUMMARY: The package HMMGEP performs cluster analysis on gene expression data using hidden Markov models. AVAILABILITY: HMMGEP, including the source code, documentation and sample data files, is available at http://www.bioinfo.tsinghua.edu.cn:8080/~rich/hmmgep_download/index.html.

Algorithms↗

Prediction of protein subcellular locations using fuzzy k-NN method.

MOTIVATION: Protein localization data are a valuable information resource helpful in elucidating protein functions. It is highly desirable to predict a protein's subcellular locations automatically from its sequence. RESULTS: In this paper, fuzzy k-nearest neighbors (k-NN) algorithm has been introduced to predict proteins' subcellular locations from their dipeptide composition. The prediction is performed with a new data set derived from version 41.0 SWISS-PROT databank, the overall predictive accuracy about 80% has been achieved in a jackknife test. The result demonstrates the applicability of this relative simple method and possible improvement of prediction accuracy for the protein subcellular locations. We also applied this method to annotate six entirely sequenced proteomes, namely Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster, Oryza sativa, Arabidopsis thaliana and a subset of all human proteins. AVAILABILITY: Supplementary information and subcellular location annotations for eukaryotes are available at http://166.111.30.65/hying/fuzzy_loc.htm

Algorithms↗

Classifying G-protein coupled receptors with bagging classification tree.

G-protein coupled receptors (GPCRs) play a key role in different biological processes, such as regulation of growth, death and metabolism of cells. They are major therapeutic targets of numerous prescribed drugs. However, the ligand specificity of many receptors is unknown and there is little structural information available. Bioinformatics may offer one approach to bridge the gap between sequence data and functional knowledge of a receptor. In this paper, we use a bagging classification tree algorithm to predict the type of the receptor based on its amino acid composition. The prediction is performed for GPCR at the sub-family and sub-sub-family level. In a cross-validation test, we achieved an overall predictive accuracy of 91.1% for GPCR sub-family classification, and 82.4% for sub-sub-family classification. These results demonstrate the applicability of this relative simple method and its potential for improving prediction accuracy.

Algorithms↗

Comparative analysis of amino acid usage and protein length distribution between alternatively and non-alternatively spliced genes across six eukaryotic genomes.

Alternative splicing has been discovered in nearly all metazoan organisms as a mechanism to increase the diversity of gene products. However, the origin and evolution of alternatively spliced genes are still poorly understood. To understand the mechanisms for the evolution of alternatively spliced genes, it may be important to study the differences between alternatively and non-alternatively spliced genes. The aim of this research was to compare amino acid usage and protein length distribution between alternatively and non-alternatively spliced genes across six nearly complete eukaryotic genomes, including those of human (Homo sapiens), mouse (Mus musculus), rat (Rattus norvegicus), fruit fly (Drosophila melanogaster), Caenorhabditis elegans, and bovine (Bos taurus). Our results have suggested the following: (1) across the six species, alternatively and non-alternatively spliced genes have very similar tendency for amino acids usage for not only the overall scale but also those highly expressed genes, with all of the highly expressed genes having preferred amino acids including A, E, G, K, L, P, S, V, R, T, and D. (2) For not only the overall genes but also those highly expressed ones, the average length of the protein products of alternatively spliced genes is significantly greater than that of non-alternatively spliced ones. In contrast, distributions of protein lengths for the two groups of genes are very similar among all six species. Based on these results, we propose that alternatively spliced genes may have originated from non-alternatively spliced ones through events such as DNA mutations or gene fusion.

Alternative Splicing↗

dbNEI: a specific database for neuro-endocrine-immune interactions.

OBJECTIVES: To construct a specific database for the neuro-endocrine-immune (NEI) interactions. METHODS/RESULTS: Version 1.0 database for neuro-endocrine-immune (dbNEI) serves as a web-based knowledge resource specific for the NEI systems. dbNEI collects 1,058 NEI related signal molecules, their 940 interactions and 72 affiliated tissues from the Cell Signaling Networks database, manually selects 982 NEI papers from PubMed, and gives links to 27,848 NEI generally related genes from UniGene database. NEI related information, such as signal transductions, regulations and control subunits, is integrated. Especially, dbNEI represents as graphic visualization, by which control subunits can be automatically obtained according to the inquiring issues, the combinative queries and the NEI related diseases respectively. CONCLUSIONS: dbNEI, which can be accessed at http://bioinfo.au.tsinghua.edu.cn/dbNEIweb/, provides a knowledge environment for understanding the main regulatory systems of NEI in a molecular level.

Animals↗