PubMed Health⌕ Search

Biomedical subjects

Yan-Da Li

Publications and source records attributed to Yan-Da Li.

8 recordsLinked to original sources

[Effect of stomatin-like protein 2 (SLP-2) gene on growth and proliferation of esophageal squamous carcinoma cell line TE12].

BACKGROUND & OBJECTIVE: Stomatin-like protein 2 (SLP-2), a differentially expressed gene obtained from esophageal squamous cell carcinoma (ESCC) matched tissues using cDNA microarray, is over-expressed in ESCC tissues. This study was to confirm over-expression of SLP-2 in ESCC tissues, to construct eukaryotic expression plasmid of SLP-2, and investigate the role of up-regulation of SLP-2 in initiation and progression of ESCC. METHODS: Semi-quantitative reverse transcription-polymerase chain reaction (RT-PCR) was used to detect expression of SLP-2 in ESCC tissues. Full-length open-reading frame (ORF) of SLP-2 was amplified from cDNA of normal esophageal epithelia by PCR, and inserted into pcDNA3.1/myc-His(-) to construct a eukaryotic expression plasmid of SLP-2. Then the recombinant was transfected into ESCC cell line TE12. Positive cell clones were selected by semi-quantitative RT-PCR. MTT assay, plate clone formation assay, flow cytometry, and MTS assay were performed to measure the effect of SLP-2 on TE12 cells. RESULTS: SLP-2 was over-expressed in ESCC tissues. Sense/antisense eukaryotic expression plasmids of SLP-2 were constructed. Antisense transfection of SLP-2 gene led to S phase arrest, decreased expression of SLP-2 in TE12 cells, and suppressed cell growth and proliferation, and cell adhesive ability. CONCLUSION: Up-regulation of SLP-2 gene might contribute to hyperproliferation of TE12 cells, and metastasis of ESCC.

Blood Proteins↗

VSD: a database for schizophrenia candidate genes focusing on variations.

Schizophrenia is a common mental disease characterized by delusions, hallucinations, and formal thought disorder. It has been demonstrated with genetic evidence that the disease is a polygenic disorder. Pharmacological, neurochemical, and clinical studies have suggested a number of schizophrenia susceptibility loci. In order to systematically search for genes with small effect in the development of schizophrenia, a database called VSD was established to provide variation data for publicly available candidate genes. Most of the genes encode neurotransmitter receptors, neurotransmitter transporters, and the enzymes involved in their metabolism. Other candidate genes extracted from published literature are also included. The variation information has been collected from publicly available mutation and polymorphism databases such as dbSNP, HGVbase, and OMIM, with single nucleotide polymorphism (SNP) being the most abundant form of collected variations. Reference sequences from NCBI's RefSeq database are used as references when positioning variation at transcript and protein levels. The nonsynonymous SNPs (nsSNPs) that lead to amino acid changes in the functional sites or domains of proteins are distinguished since they are more likely to affect protein function and would be target SNPs for association studies. In addition to variation data, gene descriptions, enzyme information, and other biological information for each gene locus are also included. The latest version of VSD contains 23,648 variations assigned to a total of 186 genes. Five-hundred eighty-eight domains and sites annotated in the SWISS-PROT and InterPro databases are found to contain nsSNPs. VSD may be accessed via the World Wide Web (www.chgb.org.cn/vsd.htm) and will be developed as an up-to-date and comprehensive locus-specific resource for identifying susceptibility genes for schizophrenia.

Databases, Nucleic Acid↗

[Analysis, identification and correction of some errors of model refseqs appeared in NCBI Human Gene Database by in silico cloning and experimental verification of novel human genes].

We found that human genome coding regions annotated by computers have different kinds of many errors in public domain through homologous BLAST of our cloned genes in non-redundant (nr) database, including insertions, deletions or mutations of one base pair or a segment in sequences at the cDNA level, or different permutation and combination of these errors. Basically, we use the three means for validating and identifying some errors of the model genes appeared in NCBI GENOME ANNOTATION PROJECT REFSEQS: (I) Evaluating the support degree of human EST clustering and draft human genome BLAST. (2) Preparation of chromosomal mapping of our verified genes and analysis of genomic organization of the genes. All of the exon/intron boundaries should be consistent with the GT/AG rule, and consensuses surrounding the splice boundaries should be found as well. (3) Experimental verification by RT-PCR of the in silico cloning genes and further by cDNA sequencing. And then we use the three means as reference: (1) Web searching or in silico cloning of the genes of different species, especially mouse and rat homologous genes, and thus judging the gene existence by ontology. (2) By using the released genes in public domain as standard, which should be highly homologous to our verified genes, especially the released human genes appeared in NCBI GENOME ANNOTATION PROJECT REFSEQS, we try to clone each a highly homologous complete gene similar to the released genes in public domain according to the strategy we developed in this paper. If we can not get it, our verified gene may be correct and the released gene in public domain may be wrong. (3) To find more evidence, we verified our cloned genes by RT-PCR or hybrid technique. Here we list some errors we found from NCBI GENOME ANNOTATION PROJECT REFSEQs: (1) Insert a base in the ORF by mistake which causes the frame shift of the coding amino acid. In detail, abase in the ORF of a gene is a redundant insertion, which causes a reading frame shift in the translation of an alternative protein, such as LOC124919 is wrong form of C17 orf32 (with mouse and rat orthologs determined by us). (2) Put together by mistake (with force). This is a wrong assembly of non-relating cDNA segment, such as LOC147007 is wrong form of C17orf32. (3) Mistakenly insert a base or one section of cDNA in the ORF which causes it ending beforehand, only coding cDNA sequence of N-terminal amino acids, incomplete. For example, LOC123722 is wrong form of SPRYD1, and even the human hypothetical gene LOC126250 or PDCD5 is wrong form of our PDCD5 (TFAR19). (4) Incomplete, only coding cDNA sequence of C-terminal amino acids. For example, human LOC149076 and mouse LOC230761 are wrong form of our verified human ZNF362 and mouse Zfp362, respectively. (5) Incomplete, only coding one section of coding protein cDNA sequence of correct gene ORF, lacking N-terminal and C-terminal amino acids sequence, and at the same time, mistakenly anticipates the first non-initiation codon amino acid of the incomplete protein amino acid as the initiation codon, e.g. anticipating L as M. For example, LOC200084 is wrong form of ZNF362. (6) Mistakenly insert a base or one section of cDNA in the ORF, wrongly causing unwanted termination codon before the insertion, so the coding protein lacks the first part of the amino acids. For example, the GenBank Acc. No. AL096883 ( LOCUS No. HS323M22B) is wrong form of an experimentally verified human NM_012263 with mouse ortholog of BC010510 determined. (7) It may regard the polluted genomic sequence as complete gene cDNA sequence and anticipate the so-called single exon gene, even the real one, only a small ORF in the very long single exon mRNA, while there really exists termination code in the same phase of the upper part of the ORF initiation code, no other characters accord with the gene's condition. For example, LOC91126 is wrong form of ZNF362. (8) The anticipated genes only have ORF which has no EST proofs on both terminal sides. Depending on this ORF, a complete gene cDNA with double support of EST and human genome (there are termination codes at the same phase of the upper part of ORF) which indicates the anticipated ORF reference sequence may be incorrect. For example, LOC164395 may be wrong form of novel human gene bankit4590055. (9) A similar but smaller protein-coding gene is anticipated in the range of the human genome sequence that has the support of EST experimental proof, so other new anticipated gene may be incorrect. For example, LOC167563 may be wrong form of CMYA5. However,these errors can be corrected or avoided by using our strategy. Here we give one example in detail: Comparision of the sequence SPRYD1 with human hypothetical gene LOC123722. The TAA bases in the position of 478-480 in LOC123722 cDNA is redundant, which causes a reading frame shift in the translation of an alternative protein. The redundancy of GTAAA of LOC123722 is not supported by our experimental clone,and is almost fully rejected by human EST alignment, and is shown as the next intron sequence by genomic GT/AG organization analysis. The verification of cDNA or genomic DNA sequence of SPRYD1 implies that LOC123722 has a wrong stop codon within its ORF because of the prediction program, thus being not complete cds. To sum up, by combining bioinformatics analyses with experimental verification, we have found that there are many errors of at least nine kinds appeared in NCBI GENOME ANNOTATION PROJECT REFSEQs through BLAST of our cloned genes in non-redundant database, and our strategy is helpful in correcting them, such as LOC14907, LOC200084 and LOC91126 (all of them should be ZNF362, but are three different kinds of wrong forms of ZNF362), three model reference sequences predicted from NCBI contig NT_004511 by automated computational analysis using gene prediction method, or such as LOC124919 and LOC147007 (both should be C17orf32, but are two different kinds of wrong forms of C17orf32), two model reference sequences predicted from NCBI contig NT_010808 by automated computational analysis using gene prediction method. Therefore, the correct identification and annotation of novel human genes may be still a heavy task, which can be finished within a long period of time. So human genome coding regions annotated by computer should be used with caution. The articles published in the past did not clearly point out the existence of mistakes in the NCBI human gene mode reference sequence. At the Seventh International Human Genome Conference held in April 2002, we first published the researching result on this aspect in the communication form of Posterly insert a base or one section of cDNA in the ORF, wrongly causing unwanted termination codon before the insertion, so the coding protein lacks the first part of the amino acids. For example, the GenBank Acc. No. AL096883 ( LOCUS No. HS323M22B) is wrong form of an experimentally verified human NM_012263 with mouse ortholog of BC010510 determined. (7) It may regard the polluted genomic sequence as complete gene cDNA sequence and anticipate the so-called single exon gene, even the real one, only a small ORF in the very long single exon mRNA, while there really exists termination code in the same phase of the upper part of the ORF initiation code, no other characters accord with the gene's condition. For example, LOC91126 is wrong form of ZNF362. (8) The anticipated genes only have ORF which has no EST proofs on both terminal sides. Depending on this ORF, a complete gene cDNA with double support of EST and human genome (there are termination codes at the same phase of the upper part of ORF) which indicates the anticipated ORF reference sequence may be incorrect. For example, LOC164395 may be wrong form of novel human gene bankit4590055. (9) A similar but smaller protein-coding gene is anticipated in the range of the human genome sequence that has the support of EST experimental proof, so other new anticipated gene may be incorrect. For example, LOC167563 may be wrong form of CMYA5. However, these errors can be corrected or avoided by using our strategy. Here we give one example in detail: Comparision of the sequence SPRYD1 with human hypothetical gene LOC123722. The TAA bases in the position of 478-480 in LOC123722 cDNA is redundant, which causes a reading frame shift in the translation of an alternative protein. The redundancy of GTAAA of LOC123722 is not supported by our experimental clone, and is almost fully rejected by human EST alignment, and is shown as the next intron sequence by genomic GT/AG organization analysis. The verification of cDNA or genomic DNA sequence of SPRYD1 implies that LOC123722 has a wrong stop codon within its ORF because of the prediction program, thus being not complete cds. To sum up, by combining bioinformatics analyses with experimental verification, we have found that there are many errors of at least nine kinds appeared in NCBI GENOME ANNOTATION PROJECT REFSEQs through BLAST of our cloned genes in non-redundant database, and our strategy is helpful in correcting them, such as LOC14907, LOC200084 and LOC91126 (all of them should be ZNF362, but are three different kinds of wrong forms of ZNF362), three model reference sequences predicted from NCBI contig NT_004511 by automated computational analysis using gene prediction method, or such as LOC124919 and LOC147007 (both should be C17orf32, but are two different kinds of wrong forms of C17orf32), two model reference sequences predicted from NCBI contig NT_010808 by automated computational analysis using gene prediction method. Therefore, the correct identification and annotation of novel human genes may be still a heavy task, which can be finished within a long period of time. So human genome coding regions annotated by computer should be used with caution. (ABSTRACT TRUNCATED)

Amino Acid Sequence↗

[Correction of five different types of errors of model REFSEQs appeared in NCBI human gene database only by using two novel human genes C17orf32 and ZNF362].

Found that there exist many mistakes in the REFSEQ issued in the genome annotation project of NCBI, the result of which indicates that people be cautious in using REFSEQ database in NCBI. By adopting the technical route combining bioinformatics analysis and experimental verification, through the comparison of the cloned genes in the non-redundant database, we found that there were many mistakes in the computer annotation human genome coding sequences that were issued on the internet. First we quoted nine wrong types of novel human genes anticipated by NCBI GENOME Annotation Project. Here we give one example in detail: (1) Comparison of the sequences between novel human gene C17orf32 and hypothetical human gene LOC124919. LOC123722 is a modified sequence of C17orf32 cDNA with an inserted G between 406 -407 nucleotides. The base G in the 401 position of LOC123722 cDNA is a redundant insert, which causes a reading frame shift in the translation of an alternative protein. This inserted G has not been found in our experimental clone, and is fully rejected by human EST alignment, and is shown as a redundance by genomic GT/AG organization analysis. (2) Comparison of the sequences between novel human gene C17orf32 and hypothetical human gene LOC147007. C17orf32 gene (ORF from 31 to 657 nucleotides) is located on human chromosome 17(Accession No. NT_010808.7), and is only linked with a hypothetical human gene LOC147007 (ORF from 55 to 435 nucleotides) at present. This hypothetical human gene sequence has not been verified by experiment, and is a wrong form of our verified C17orf32 gene. The full-length 1 679 bp cDNA sequence of C17orf32 exhibits overall homology to that of LOC147007 of 625 bp mRNA, with matching percentage of 37% in 36% of total window over the full-length nucleotide, especially 121 approximately 366 bp of LOC147007 is just the same as 316 approximately 561 bp of C17orf32. Thus, the 126 aa protein encoded by XP_097165 of LOC147007 exhibits overall homology to the 208 aa protein encoded by C17orf32, with matching percentage of 50% in 48% of total window over the full-length protein, especially 23 approximately 104 aa of XP_097165 is just the same as 96 approximately 177 aa of C17orf32 protein. Both flanking regions of LOC147007 outside the same ORF central part are wrong assembly of non-relative cDNA. In addition, we have in silico cloned a novel mouse gene, ORF32 (open reading frame 32) with TPA accession number of BK000258, which is the mouse ortholog of human C17orf32. Our strategy is helpful in both finding out more novel human genes and correcting the mistakes in the REFSEQs issued by NCBI genome annnotation project. For example, we adopted the gene anticipating method, through automatic calculation and analysis, anticipated two modes reference sequences (LOC124919 and LOC147007) from NCBI contig NT_ 010808. Both of them should be C17orf32, but the fact is that both of them are various wrong forms of C17orf32, respectively are the first type and second type of mistakes. Another example, we adopted gene anticipation method, through automatic calculation and analysis, anticipated three modes reference sequences (LOC14907, LOC200084 and LOC91126) from NCBI contig NT_004511 which really are one type of gene of ZNF362, but submitted three different wrong forms of ZNF362, respectively are: the fourth, fifth, and seventh type of mistakes. We can correct or avoid the currently wrong human genome coding sequence by using in silico clone and combining experimental verification. People should be cautious in treating the computer's annotation which may exist all type of wrong human genome coding sequences. The correct identification and annotation of the novel human genes still remain to be a long and arduous task.

Amino Acid Sequence↗

[The application of human mutation databases].

Researches on genome mutation are becoming more and more important with the finish of human genome DNA draft. This review is to classify the existing human mutation databases, including mutation database, SNP(single nucleotide polymorphisms) databases, mutation databases about disease, mutation databases about proteins, mutation databases about map and mutation information about specific gene. We also give advice on how to utilize these mutation databases, and discuss problems of existing databases.

Databases, Factual↗

Identifying splicing sites in eukaryotic RNA: support vector machine approach.

We introduce a new method for splicing sites prediction based on the theory of support vector machines (SVM). The SVM represents a new approach to supervised pattern classification and has been successfully applied to a wide range of pattern recognition problems. In the process of splicing sites prediction, the statistical information of RNA secondary structure in the vicinity of splice sites, e.g. donor and acceptor sites, is introduced in order to compare recognition ratio of true positive and true negative. From the results of comparison, addition of structural information has brought no significant benefit for the recognition of splice sites and had even lowered the rate of recognition. Our results suggest that, through three cross validation, the SVM method can achieve a good performance for splice sites identification.

Algorithms↗

Suppressive effects of a Chinese herbal medicine qing-luo-yin extract on the angiogenesis of collagen-induced arthritis in rats.

Qing-Luo-Yin (QLY), a traditional Chinese herbal medicine for the treatment of rheumatoid arthritis, is a combination of the extracts of Sophora flavescens Ait., Phellodendron amurense Rupr., Sinomenium acutum Rehd. et Wils. and Dioscorea hypoglauca Palib. The suppressive effect of QLY on the development of angiogenesis was investigated in a rat model of collagen-induced arthritis (CIA). QLY (0.3 g/kg) was orally administered daily for 27 days. Neo-angiogenesis, pannus and cartilage damage, the expression of metalloproteinases (MMP)-3 messenger RNA (mRNA) and the level of the tissue inhibitor of matrix metalloproteinase (TIMP)-1 in the synovium were examined by histology, in situ hybridization and immunohistochemiscal assays, respectively. It was observed that the articular morphological alterations, the over-expression of MMP-3 mRNA and the reduced production of TIMP-1 in CIA rats were significantly ameliorated by QLY. QLY performed about as effectively as tripterygium glycosidorum tablets (0.1 g/kg) extracted from Tripterygium wilfordii Hook. f.. These results indicate that QLY exerts a suppressive effect on the angiogenesis of CIA rats, and suggest that the therapeutic effect of QLY could be due to restoring the balance of MMP-3 and TIMP-1 in rheumatoid synovium.

Animals↗

Anti-Helicobacter pylori immunoglobulin G (IgG) and IgA antibody responses and the value of clinical presentations in diagnosis of H. pylori infection in patients with precancerous lesions.

AIM: To determine the prevalence of Helicobacter pylori (H. pylori) infection, the serum anti-H. pylori immunoglobulin G (IgG) and IgA antibody responses, and the value of clinical presentations in diagnosis of H. pylori infection in patients with gastric atrophy, intestinal metaplasia and dysplasia. METHODS: H. pylori infection was detected by histology in 209 patients with mild chronic atrophic gastritis (CAG, n=76), severe CAG (n=22), mild intestinal metaplasia (IM, n=22), severe IM (n=58), or dysplasia (DYS, n=31). Serum anti-H. pylori IgG and IgA were double sampled and evaluated by enzyme-linked immunoadsordent assays. 35 clinical presentations were observed and their relationship with H. pylori infection was analyzed by the k-means cluster method. RESULTS: Both IgG and IgA levels in H. pylori positive patients were significantly higher than those negative for H. pylori (P<0.001-0.01). The prevalence of H. pylori was highest in severe IM (84.5 %), and lowest in mild CAG (51.3 %) (P<0.01). They were similar in severe CAG (68.2 %), mild IM (72.7 %), and DYS (67.7 %). In H. pylori positive patients, the IgG levels in severe CAG were significantly higher than those in mild CAG (P<0.01). In H. pylori negative patients, both IgG and IgA levels increased remarkably in severe IM, compared to those in mild IM (P<0.01-0.05). H. pylori infection exhibited no association with patient's gender (62.1 % in males; 71.7 % in females) and age (r=0.0814, P=0.241). The diagnostic accuracy based on 35 clinical presentations was 65.7 %. It could be improved by 5.7 % when only the assemblage of digestive symptoms were engaged, or by 8.6 % when the pathogenic factors, general status and grossoscopy were combined. The diagnostic accuracy could be decreased when only the general symptoms were engaged, or when the pathogenic factors were accompanied with some common digestive symptoms. CONCLUSION: H. pylori infection is a major risk factor for the process from atrophy, IM to DYS of gastric mucosa. Serum IgG and IgA are good indicators to evaluate this progress with a certain arrearage. Investigation on the effective assemblages of clinical presentations may provide a better understanding in the pathogenesis, diagnosis and treatment for H. pylori infection.

Aged↗