PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Expressed Sequence Tags”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Potato expressed sequence tag generation and analysis using standard and unique cDNA libraries.

To help develop an understanding of the genes that govern the developmental characteristics of the potato (Solanum tuberosum), as well as the genes associated with responses to specified pathogens and storage conditions, The Canadian Potato Genome Project (CPGP) carried out 5' end sequencing of regular, normalized and full-length cDNA libraries of the Shepody potato cultivar, generating over 66,600 expressed sequence tags (ESTs). Libraries sequenced represented tuber developmental stages, pathogen-challenged tubers, as well as leaf, floral developmental stages, suspension cultured cells and roots. All libraries analysed to date have contributed unique sequences, with the normalized libraries high on the list. In addition, a low molecular weight library has enhanced the 3' ends of our sequence assemblies. Using the combined assembly dataset, unique tuber developmental, cold storage and pathogen-challenged sequences have been identified. A comparison of the ESTs specific to the pathogen-challenged tuber and foliar libraries revealed minimal overlap between these libraries. Mixed assemblies using over 189,000 potato EST sequences from CPGP and The Institute for Genomics Research (TIGR) has revealed common sequences, as well as CPGP- and TIGR-unique sequences.

DNA, Complementary↗

Census of genes expressed in porcine embryos and reproductive tissues by mining an expressed sequence tag database based on human genes.

A total of 98,898 expressed sequence tags (ESTs) derived from embryos and reproductive tissues in pigs were identified in the GenBank "est_others" database. Pig embryos were collected at 11, 12, 13, 14, 15, 20, 30, and 45 days after gestation. The reproductive tissues were sampled from testis, ovary, endometrium, hypothalamus, anterior pituitary, uterus, and placenta. A gene-oriented approach was developed to annotate these porcine EST sequences to census the genes expressed from these sources. Of the 33 308 mRNA sequences from the human genes used as references (data accessed on 1 November 2002), 9410 had the porcine EST homologs expressed in embryos and 11 795 had the EST homologs expressed in reproductive tissues. The entire genome contributes at least 28.3% of its genes to embryo development and 35.4% of its genes to reproduction. Using the EST entry numbers as indicators of gene expression, we determined that the gene expression patterns differ significantly between embryos and reproductive tissues in pigs. The basic active genes were identified for each source, but most of them are not coexpressed abundantly. Few genes were expressed on the Y chromosome (P < 0.01), but they may represent counterparts of the double-dose genes that remain active in an inactivated X chromosome in females but are needed for proper development and growth. The census provides a panel of transcripts in a broad sense that can be used as targets to study the mechanisms involved in embryo development and reproduction in pigs and other mammals, including humans.

Animals↗

Identification and analysis of Arabidopsis expressed sequence tags characteristic of non-coding RNAs.

Sequencing of the Arabidopsis genome has led to the identification of thousands of new putative genes based on the predicted proteins they encode. Genes encoding tRNAs, ribosomal RNAs, and small nucleolar RNAs have also been annotated; however, a potentially important class of genes has largely escaped previous annotation efforts. These genes correspond to RNAs that lack significant open reading frames and encode RNA as their final product. Accumulating evidence indicates that such "non-coding RNAs" (ncRNAs) can play critical roles in a wide range of cellular processes, including chromosomal silencing, transcriptional regulation, developmental control, and responses to stress. Approximately 15 putative Arabidopsis ncRNAs have been reported in the literature or have been annotated. Although several have homologs in other plant species, all appear to be plant specific, with the exception of signal recognition particle RNA. Conversely, none of the ncRNAs reported from yeast or animal systems have homologs in Arabidopsis or other plants. To identify additional genes that are likely to encode ncRNAs, we used computational tools to filter protein-coding genes from genes corresponding to 20,000 expressed sequence tag clones. Using this strategy, we identified 19 clones with characteristics of ncRNAs, nine putative peptide-coding RNAs with open reading frames smaller than 100 amino acids, and 11 that could not be differentiated between the two categories. Again, none of these clones had homologs outside the plant kingdom, suggesting that most Arabidopsis ncRNAs are likely plant specific. These data indicate that ncRNAs represent a significant and underdeveloped aspect of Arabidopsis genomics that deserves further study.

Algorithms↗

GrainGenes 2.0. an improved resource for the small-grains community.

GrainGenes (http://wheat.pw.usda.gov) is an international database for genetic and genomic information about Triticeae species (wheat [Triticum aestivum], barley [Hordeum vulgare], rye [Secale cereale], and their wild relatives) and oat (Avena sativa) and its wild relatives. A major strength of the GrainGenes project is the interaction of the curators with database users in the research community, placing GrainGenes as both a data repository and information hub. The primary intensively curated data classes are genetic and physical maps, probes used for mapping, classical genes, quantitative trait loci, and contact information for Triticeae and oat scientists. Curation of these classes involves important contributions from the GrainGenes community, both as primary data sources and reviewers of published data. Other partially automated data classes include literature references, sequences, and links to other databases. Beyond the GrainGenes database per se, the Web site incorporates other more specific databases, informational topics, and downloadable files. For example, unique BLAST datasets of sequences applicable to Triticeae research include mapped wheat expressed sequence tags, expressed sequence tag-derived simple sequence repeats, and repetitive sequences. In 2004, the GrainGenes project migrated from the AceDB database and separate Web site to an integrated relational database and Internet resource, a major step forward in database delivery. The process of this migration and its impacts on database curation and maintenance are described, and a perspective on how a genomic database can expedite research and crop improvement is provided.

Breeding↗

In silico identification and expression of SLC30 family genes: an expressed sequence tag data mining strategy for the characterization of zinc transporters' tissue expression.

BACKGROUND: Intracellular zinc concentration and localization are strictly regulated by two main protein components, metallothioneins and membrane transporters. In mammalian cells, two membrane transporters family are involved in intracellular zinc homeostasis: the uptake transporters called SLC39 or Zip family and the efflux transporters called SLC30 or ZnT family. ZnT proteins are members of the cation diffusion facilitator (CDF) family of metal ion transporters. RESULTS: From genomic databanks analysis, we identified the full-length sequences of two novel SLC30 genes, SLC30A8 and SLC30A10, extending the SLC30 family to ten members. We used an expressed sequence tag (EST) data mining strategy to determine the pattern of ZnT genes expression in tissues. In silico results obtained for already studied ZnT sequences were compared to experimental data, previously published. We determined an overall good correlation with expression pattern obtained by RT-PCR or immunomethods, particularly for highly tissue specific genes. CONCLUSION: The method presented herein provides a useful tool to complete gene families from sequencing programs and to produce preliminary expression data to select the proper biological samples for laboratory experimentation.

Amino Acid Sequence↗

Adult midgut expressed sequence tags from the tsetse fly Glossina morsitans morsitans and expression analysis of putative immune response genes.

BACKGROUND: Tsetse flies transmit African trypanosomiasis leading to half a million cases annually. Trypanosomiasis in animals (nagana) remains a massive brake on African agricultural development. While trypanosome biology is widely studied, knowledge of tsetse flies is very limited, particularly at the molecular level. This is a serious impediment to investigations of tsetse-trypanosome interactions. We have undertaken an expressed sequence tag (EST) project on the adult tsetse midgut, the major organ system for establishment and early development of trypanosomes. RESULTS: A total of 21,427 ESTs were produced from the midgut of adult Glossina morsitans morsitans and grouped into 8,876 clusters or singletons potentially representing unique genes. Putative functions were ascribed to 4,035 of these by homology. Of these, a remarkable 3,884 had their most significant matches in the Drosophila protein database. We selected 68 genes with putative immune-related functions, macroarrayed them and determined their expression profiles following bacterial or trypanosome challenge. In both infections many genes are downregulated, suggesting a malaise response in the midgut. Trypanosome and bacterial challenge result in upregulation of different genes, suggesting that different recognition pathways are involved in the two responses. The most notable block of genes upregulated in response to trypanosome challenge are a series of Toll and Imd genes and a series of genes involved in oxidative stress responses. CONCLUSIONS: The project increases the number of known Glossina genes by two orders of magnitude. Identification of putative immunity genes and their preliminary characterization provides a resource for the experimental dissection of tsetse-trypanosome interactions.

Aging↗

Using expressed sequence tag databases to identify ovarian genes of interest.

GenBank contains 4879 expressed sequence tags (EST) derived from four non-normalized human ovarian cDNA libraries. Of these EST, 2646 are contributors to UniGene clusters and have UniGene numbers. The EST map to 1206 distinct UniGenes. A gene expression profile was established for the human ovary by identifying the abundance of each UniGene cluster and its corresponding annotation. The most highly expressed transcripts were for proteins associated with protein synthesis (ribosomal proteins, elongation factors, thymosins, etc.). However, there are also transcripts for genes of unknown function that are ovary-specific. This ovarian gene expression profile provides useful data for the design of DNA microarrays targeted at ovarian function and highlights novel sequences that warrant further investigation.

Databases, Nucleic Acid↗

EbEST: an automated tool using expressed sequence tags to delineate gene structure.

Large numbers of expressed sequence tags (ESTs) continue to fill public and private databases with partial cDNA sequences. However, using this huge amount of ESTs to facilitate gene finding in genomic sequence imposes a challenge, especially to wet-lab scientists who often have limited computing resources. In an effort to consolidate the information hidden in the vast number of ESTs into a readable and manageable format, we have developed EbEST-a program that automates the process of using ESTs to help delineate gene structure in long stretches of genomic sequence. The EbEST program consists of three functional modules-the first module separates homologous ESTs into clusters and identifies the most informative ESTs within each cluster; the second module uses the informative ESTs to perform gapped alignment and to predict the exon-intron boundary; and the third module generates text file and graphic outputs that illustrate the orientation, exonic structure, and untranslated regions (UTRs) of putative genes in the genomic sequence being analyzed. Evaluation of EbEST with 176 human genes from the ALLSEQ set indicated that it performed in-line with several existing gene finding programs, but was more tolerant to sequencing errors. Furthermore, when EbEST was challenged with query sequences that harbor more than one gene, it suffered only a slight drop in performance, whereas the performance of the other programs evaluated decreased more. EbEST may be used as a stand-alone tool to annotate human genomic sequences with EST-derived gene elements, or can be used in conjunction with computational gene-recognition programs to increase the accuracy of gene prediction. [EbBEST is available at http://EbEST.ifrc.mcw.edu]

Base Sequence↗

In silico identification of components of the Toll-like receptor (TLR) signaling pathway in clustered chicken expressed sequence tags (ESTs).

We have described a bioinformatic approach that involves the clustering of expressed sequence tags (ESTs) to reveal homologs of the Toll-like receptor (TLR) pathway in the chicken. Homology searching of proteins, predicted to be encoded by these EST clusters, resulted in the in silico identification of full-length sequences for Toll-interacting protein (Tollip), IL-1 receptor-associated kinase 4 (IRAK-4), myeloid differentiation factor 88 adapter-like (Mal), TGF beta-activated kinase 1 binding protein 1 (TAB1). We also determined partial sequence information for myeloid differentiation factor 88 (MyD88), two novel TLRs, TNF receptor-associated factor 6 (TRAF6), TGF beta-activated kinase 1 (TAK1), TAB2, inhibitor of nuclear factor kappa B kinase alpha (IKK alpha) and IKK beta. This bioinformatics study has confirmed the evolutionary conservation of the TLR pathway in chicken and demonstrated its essential homology to the TLR pathway in mammals. We have identified in silico the full-length sequence for liver-expressed antimicrobial peptide 2 (LEAP-2). This is the first time a non-mammalian LEAP-2 has been described.

Amino Acid Sequence↗

A survey of genes expressed in mouse embryonal carcinoma F9 cells: characterization of expressed sequence tags matching no known genes.

We prepared 2,132 expressed sequence tags (ESTs) from undifferentiated mouse embryonal carcinoma F9 cells and found that 1,416 match known gene and/or protein sequences [Nishiguchi et al. (1996) J. Biochem. 119, 749-767]. To obtain information on the functions of the remaining 716 unidentified ESTs and to develop a system for characterizing ESTs matching no known genes, we analyzed their sequences by (i) repeated database searches, using the BLASTN, BLASTX, TBLASTX, and FASTA programs, (ii) using computer programs developed or modified for this work, such as the WFASTA, ORFTRNS, and MFASTA programs, together with the DBPROSITE and GRAIL programs, and (iii) examining the expression patterns of the corresponding mRNAs in F9 cells and several organs of adult mice, using the digoxigenin-labeled dot-blot method. We found that 216 of the 716 ESTs match known gene and/or protein sequences, and 307 show significant similarities to these sequences, with a Poisson p-value < 0.01. The strategy and usefulness of such analysis for characterizing unidentified ESTs are discussed.

Animals↗

Tissue-Specific Expressed Sequence Tags from the Black Tiger Shrimp Penaeus monodon.

: Expressed sequence tag data were generated from complementary DNA libraries created from cephalothorax, eyestalk, and pleopod tissue of the black tiger shrimp (Penaeus monodon). Significant database matches were found for 48 of 83 nuclear genes sequenced from the cephalothorax library, 22 of 55 nuclear genes from the eyestalk library, and 6 of 13 nuclear genes from the pleopod library. The putative identities of these genes reflected the expected tissue specificity. For example, genes for digestive enzymes were identified from the cephalothorax library and genes involved in the visual and neuroendocrine system from the eyestalk library. A few sequences matched anonymous EST or genomic sequences, and others contained mini-satellite or microsatellite repeat sequences. The remainder, 31 from the cephalothorax library, 25 from the eyestalk library, and 5 from the pleopod library, were sequences of high nucleotide complexity with no matches in any database searched and thus may represent novel genes.

Journal Article↗

Analysis of expressed sequence tags (ESTs) in the ciliated protozoan Tetrahymena thermophila.

To assess the utility of expressed sequence tag (EST) sequencing as a method of gene discovery in the ciliated protozoan Tetrahymena thermophila, we have sequenced either the 5' or 3' ends of 157 clones chosen at random from two cDNA libraries constructed from the mRNA of vegetatively growing cultures. Of 116 total non-redundant clones, 8.6% represented genes previously cloned in Tetrahymena. Fifty-two percent had significant identity to genes from other organisms represented in GenBank, of which 92% matched human proteins. Intriguing matches include an opioid-regulated protein, a glutamate-binding protein for an NMDA-receptor, and a stem-cell maintenance protein. Eleven-percent of the non-Tetrahymena specific matches were to genes present in humans and other mammals but not found in other model unicellular eukaryotes, including the completely sequenced Saccharomyces cerevisiae. Our data reinforce the fact that Tetrahymena is an excellent unicellular model system for studying many aspects of animal biology and is poised to become an important model system for genome-scale gene discovery and functional analysis.

Animals↗

Shotgun sequencing of the human transcriptome with ORF expressed sequence tags.

Theoretical considerations predict that amplification of expressed gene transcripts by reverse transcription-PCR using arbitrarily chosen primers will result in the preferential amplification of the central portion of the transcript. Systematic, high-throughput sequencing of such products would result in an expressed sequence tag (EST) database consisting of central, generally coding regions of expressed genes. Such a database would add significant value to existing public EST databases, which consist mostly of sequences derived from the extremities of cDNAs, and facilitate the construction of contigs of transcript sequences. We tested our predictions, creating a database of 10,000 sequences from human breast tumors. The data confirmed the central distribution of the sequences, the significant normalization of the sequence population, the frequent extension of contigs composed of existing human ESTs, and the identification of a series of potentially important homologues of known genes. This approach should make a significant contribution to the early identification of important human genes, the deciphering of the draft human genome sequence currently being compiled, and the shotgun sequencing of the human transcriptome.

Animals↗

Differential gene expression of rat neonatal heart analyzed by suppression subtractive hybridization and expressed sequence tag sequencing.

Heart diseases have been one of the major killers among the human population worldwide. Because the vast majority of cardiomyocytes cannot regenerate once they cease to proliferate shortly after birth, functionally significant myocardial regeneration is not observed clinically. Whether these cells are terminally differentiated and permanently withdrawn from the cell cycle is controversial, but broadening our understanding of the rapid switch from hyperplastic to hypertrophic growth of cardiomyocytes during neonatal myocardial development may shed light on novel cardiovascular therapies. By suppression subtractive hybridization (SSH) and expressed sequence tag (EST) sequencing, we analyzed the differential gene expression of rat neonatal heart. SSH yielded subtracted and normalized cDNA libraries and enhanced the probability of detecting ESTs, which represent genes pertinent to signal transduction/cell regulation and replication/transcription/translation machinery, as compared to the traditional EST sequencing of heart cDNA libraries.

Animals↗

Expressed sequence tags of radish flower buds and characterization of a CONSTANS LIKE 1 gene.

Expressed sequence tag (EST) analysis was conducted for young flower buds of radish plants. Among a total of 66 ESTs examined, 40 showed a significant similarity to previously identified genes. Twenty-eight ESTs were similar to proteins identified in other plants, 11 were similar to eukaryotic proteins other than plants, and one was similar to a prokaryotic protein. Four clones were selected for further studies. EST clone 81, which showed a homology to germin-like proteins was expressed more abundantly in leaves and roots as compared to flower buds. Clone 105 was highly homologous to the translation inhibitor protein and was expressed in all three organs, but the expression level was higher in flower buds and roots. Another EST clone, 133, which shared a significant similarity with the Ran-binding protein, hybridized to two different size transcripts that were detectable only in flower buds. Clone 39 was a homolog of CONSTANS, which is a gene involved in controlling the flowering time in Arabidopsis. The cDNA clone of EST clone 39 containing the entire open reading frame was obtained and designated as RsCOL1 (Raphanus sativus CONSTANS LIKE 1). It was 1049 bp long and contained an open reading frame of 307 amino acid residues (calculated molecular mass = 33.1 kDa). The RsCOL1 protein contained two putative zinc finger motifs in the amino terminal region which were 59% identical to the corresponding region of the Arabidopsis CO protein. The radish protein also contained a predicted nuclear localization domain in the carboxyl terminal region which was 87% identical to the corresponding region of CO. DNA blot analysis revealed that the radish genome contained several genes similar to RsCOL1. RNA blot analysis showed that RsCOL1 was strongly expressed in flower buds at the early bolting stage, and the expression level declined as the flower bud matured. The transcript was also detectable in leaves and roots. In mature flowers, the RsCOL1 transcript was present primarily in carpels.

Amino Acid Sequence↗

Physical linkage of expressed sequence tags (ESTs) to polymorphic markers on the X chromosome.

Expressed sequence tags (ESTs) can in principle serve as specialized sequence tagged sites (STSs) to assemble a functional map of the human genome. The strategy of physically linking ESTs to the nearest genetic linkage markers should provide specific candidate genes for the X-linked diseases associated with these markers or loci. Therefore, 19 ESTs assigned to the X chromosome in the Genome Database (GDB) were analyzed. Eighteen were confirmed to be X-specific and were localized to regions of the X chromosome using a panel of somatic cell hybrids. Localization was then refined by positioning them on yeast artificial chromosome (YAC)-based maps. Seventeen ESTs identified cognate YACs by PCR screening and 12 of the ESTs have been assembled in YAC contigs containing polymorphic and other X chromosomal markers. Two of them also produced syntenically equivalent products in mouse. Thus localizing ESTs relative to polymorphic markers will help to assemble an integrated physical and transcriptional map of the chromosome and provide candidates for disease-gene searches.

Animals↗

Discovery of germ cell-specific transcripts by expressed sequence tag database analysis.

OBJECTIVE: To identify transcripts whose expression is restricted to germ cells. DESIGN: Expressed sequence tags (ESTs) from unfertilized egg libraries were utilized to perform in silico subtraction and identify germ cell-specific transcripts. SETTING: Baylor College of Medicine, Houston, Texas. ANIMAL(S): C57BL/6J/129SvEv hybrid. INTERVENTION(S): Tissue harvesting from mice. MAIN OUTCOME MEASURE(S): Identification of germ cell-specific transcripts. RESULT(S): We have used the Unigene collection of mouse cDNA libraries to identify ESTs derived from unfertilized egg libraries. A total of 3,499 ESTs were identified from Knowles Solter and Ko unfertilized egg cDNA libraries. In silico subtraction identified 258 ESTs, which were found in these unfertilized egg libraries, but not in adult mouse tissue cDNA libraries. We performed reverse transcription polymerase chain reaction (RT-PCR) on multiple adult tissues with 43 selected ESTs and found 5 of them where expression was absent in heart, lung, liver, brain, spleen, stomach, intestines, kidneys, and uterus, but restricted to ovaries and testes. Three ESTs were further analyzed, and they were exclusively localized to the oocytes by in situ hybridization. CONCLUSION: We have shown that utilization of publicly available ESTs from murine EST libraries is a simple and rapid in silico approach to the identification of transcripts preferentially expressed in germ cells.

Animals↗

Immune gene discovery by expressed sequence tags generated from hemocytes of the bacteria-challenged oyster, Crassostrea gigas.

An expressed sequence tag program was undertaken to isolate genes involved in defense mechanisms of the Pacific oyster, Crassostrea gigas. Putative function could be assigned to 54% of the 1142 sequenced cDNAs. We built a public database where all EST information are accessible through numerous search profiles (http://www.ifremer.fr/GigasBase). Based on sequence similarities we identified 20 genes that may be implicated in immune function. We investigated the expression of four of these genes during bacterial challenge of oysters. Three of them were induced in response to challenge lending support to their involvement in oyster immunity. Moreover, four other genes were highly homologous to components of the NF-kappa B signaling pathway which is involved in innate immune response in Drosophila and mammals. Altogether, our results open a new way to investigate the immune response in mollusks.

Animals↗