PubMed Health⌕ Search

Biomedical subjects

Guanglan Zhang

Publications and source records attributed to Guanglan Zhang.

7 recordsLinked to original sources

Promoter profiling and coexpression data analysis identifies 24 novel genes that are coregulated with AMPA receptor genes, GRIAs.

We identified a set of transcriptional elements that are conserved and overrepresented within the promoters of human, mouse, and rat GRIAs by comparing these promoters against a collection of 10,741 gene promoters. Cells regulate functional groups of genes by coordinating the transcriptional and/or posttranscriptional mRNA levels of interacting genes. As such, it is expected that functional groups of genes share the same transcriptional features within their promoters. We found 47 genes whose promoters contain the same combination of transcriptional elements that are overrepresented within the promoters of the GRIA gene family. Coexpressed genes may be transcriptionally coregulated, which in turn suggests that these genes may play complementary roles within a particular functional context. Using microarray expression data, we found 24 (of the 47) genes that share not only a similar promoter profile with GRIAs but also a well-correlated gene expression profile and, thus, we believe these to be coregulated with GRIAs.

Animals↗

Reverse transcriptase template switching and false alternative transcripts.

Reverse transcriptase (RT) can switch from one template to another in a homology-dependent manner. In the study of eukaryotic transcripts, this propensity of RT can produce an artificially deleted cDNA, which can be wrongly interpreted as an alternative transcript. Here, we have investigated the presence of such template-switching artifacts in cDNA databases, by scanning a collection of human splice sites (Information for the Coordinates of Exons, ICE database). We have confirmed several cases at the experimental level. Artifacts represent a significant portion of apparently spliced sequences using noncanonical splice signals but are rare in the context of the whole database. However, care should be taken in the annotation of alternative transcripts, especially when the RT used is poorly thermostable and when the putative intron is flanked by direct repeats, which are the substrate for template switching.

Alternative Splicing↗

Information for the Coordinates of Exons (ICE): a human splice sites database.

We present a comprehensive database, Information for the Coordinates of Exons (ICE), of genomic splice sites (SSs) for 10,803 human genes. ICE contains 91,846 pairs of donor acceptor sites, supported by the alignment of "full-length" human mRNAs (including transcript variants) on human genomic sequences. ICE represents the largest collection of human SSs known to date and provides a significant resource to both molecular biologists and bioinformaticians alike. A user can visualize and extract genomic sequences around SSs of the donor acceptor pairs and can also visualize the primary structure of individual genes. We list in this article the 22 most frequently found canonical and noncanonical splice sites. The top four most represented donor acceptor pairs (GT-AG, GC-AG, AT-AC, and GT-GG) accounted for 99.16% of our data set. In addition, we calculated the SS matrix models for the three most common donor acceptor pairs. The database is focused on providing SSs and surrounding sequence information, associated SS and sequence characteristics, and relation to overall transcript structure. It allows targeted search and presents evidence for the gene structure.

Computational Biology↗

FIE2: A program for the extraction of genomic DNA sequences around the start and translation initiation site of human genes.

FIE2 (5' end Information Extraction v2) is a web-based program for easy identification and extraction of nucleotide sequence around the start of genes (promoter region) and their translation initiation site (TIS). Using information provided by the National Center for Biotechnology Information's (NCBI's) LocusLink, FIE2 identifies the 5'-most end of a gene on its respective chromosome based on alignment of a selected set of mRNAs representative of the gene. FIE2 then uses currently available human genome sequence information to extract the desired sequences. The accuracy of the information extracted is therefore limited by the accuracy and completeness of the sequence annotation and sequence alignment provided by LocusLink. In addition, multiple TIS positions are also occasionally presented, for example, as a result of multiple alignments of transcript variants. One of the key criteria of FIE2 is that it should extract only the correct information or attempt no extraction at all. To date, the authors are not aware of any publicly available web-based tool that uses the human genomic sequence to extract pertinent promoter- and TIS-region information in this fashion. FIE2 is freely available at http://sdmc.lit.org.sg/FIE2.0.

Base Sequence↗

Prediction of promiscuous peptides that bind HLA class I molecules.

Promiscuous T-cell epitopes make ideal targets for vaccine development. We report here a computational system, MULTIPRED, for the prediction of peptide binding to the HLA-A2 supertype. It combines a novel representation of peptide/MHC interactions with a hidden Markov model as the prediction algorithm. MULTIPREDis both sensitive and specific, and demonstrates high accuracy of peptide-binding predictions for HLA-A*0201, *0204, and *0205 alleles, good accuracy for *0206 allele, and marginal accuracy for *0203 allele. MULTIPREDreplaces earlier requirements for individual prediction models for each HLA allelic variant and simplifies computational aspects of peptide-binding prediction. Preliminary testing indicates that MULTIPRED can predict peptide binding to HLA-A2 supertype molecules with high accuracy, including those allelic variants for which no experimental binding data are currently available.

Amino Acid Motifs↗

Dragon Promoter Finder: recognition of vertebrate RNA polymerase II promoters.

Dragon Promoter Finder (DPF) locates RNA polymerase II promoters in DNA sequences of vertebrates by predicting Transcription Start Site (TSS) positions. DPF's algorithm uses sensors for three functional regions (promoters, exons and introns) and an Artificial Neural Network (ANN). Results on a large and diverse evaluation set indicate that DPF exhibits a superior predicting ability for TSS location compared to three other promoter-finding programs.

Algorithms↗