PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Expressed Sequence Tags”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

TBestDB: a taxonomically broad database of expressed sequence tags (ESTs).

The TBestDB database contains approximately 370,000 clustered expressed sequence tag (EST) sequences from 49 organisms, covering a taxonomically broad range of poorly studied, mainly unicellular eukaryotes, and includes experimental information, consensus sequences, gene annotations and metabolic pathway predictions. Most of these ESTs have been generated by the Protist EST Program, a collaboration among six Canadian research groups. EST sequences are read from trace files up to a minimum quality cut-off, vector and linker sequence is masked, and the ESTs are clustered using phrap. The resulting consensus sequences are automatically annotated by using the AutoFACT program. The datasets are automatically checked for clustering errors due to chimerism and potential cross-contamination between organisms, and suspect data are flagged in or removed from the database. Access to data deposited in TBestDB by individual users can be restricted to those users for a limited period. With this first report on TBestDB, we open the database to the research community for free processing, annotation, interspecies comparisons and GenBank submission of EST data generated in individual laboratories. For instructions on submission to TBestDB, contact tbestdb@bch.umontreal.ca. The database can be queried at http://tbestdb.bcm.umontreal.ca/.

Animals↗

Diversity of odourant binding proteins revealed by an expressed sequence tag project on male Manduca sexta moth antennae.

A small expressed sequence tag (EST) project generating 506 ESTs from 375 cDNAs was undertaken on the antennae of male Manduca sexta moths in an effort to discover olfactory receptor proteins. We encountered several clones that encode apparent transmembrane proteins; however, none is a clear candidate for an olfactory receptor. Instead we found a greater diversity of odourant binding proteins (OBPs) than previously known in moth antennae, raising the number known for M. sexta from three to seven. Together with evidence of seventeen members of the family from the Drosophila melanogaster genome project, our results suggest that insects may have many tens of OBPs expressed in subsets of the chemosensory sensilla on their antennae. These results support a model for insect olfaction in which OBPs selectively transport and present odourants to transmembrane olfactory receptors. We also found five members of a family of shorter proteins, named sensory appendage proteins (SAPs), that might also be involved in odourant transport. This small EST project also revealed several candidate odourant degrading enzymes including three P450 cytochromes, a glutathione S-transferase and a uridine diphosphate (UDP) glucosyltransferase. Several first insect homologues of proteins known from vertebrates, the nematode Caenorhabditis elegans, yeast and bacteria were encountered, and most have now also been detected by the large D. melanogaster EST project. Only thriteen entirely novel proteins were encountered, some of which are likely to be cuticle proteins.

Amino Acid Sequence↗

Transcriptome analysis of Deinagkistrodon acutus venomous gland focusing on cellular structure and functional aspects using expressed sequence tags.

BACKGROUND: The snake venom gland is a specialized organ, which synthesizes and secretes the complex and abundant toxin proteins. Though gene expression in the snake venom gland has been extensively studied, the focus has been on the components of the venom. As far as the molecular mechanism of toxin secretion and metabolism is concerned, we still knew a little. Therefore, a fundamental question being arisen is what genes are expressed in the snake venom glands besides many toxin components? RESULTS: To examine extensively the transcripts expressed in the venom gland of Deinagkistrodon acutus and unveil the potential of its products on cellular structure and functional aspects, we generated 8696 expressed sequence tags (ESTs) from a non-normalized cDNA library. All ESTs were clustered into 3416 clusters, of which 40.16% of total ESTs belong to recognized toxin-coding sequences; 39.85% are similar to cellular transcripts; and 20.00% have no significant similarity to any known sequences. By analyzing cellular functional transcripts, we found high expression of some venom related genes and gland-specific genes, such as calglandulin EF-hand protein gene and protein disulfide isomerase gene. The transcripts of creatine kinase and NADH dehydrogenase were also identified at high level. Moreover, abundant cellular structural proteins similar to mammalian muscle tissues were also identified. The phylogenetic analysis of two snake venom toxin families of group III metalloproteinase and serine protease in suborder Colubroidea showed an early single recruitment event in the viperids evolutionary process. CONCLUSION: Gene cataloguing and profiling of the venom gland of Deinagkistrodon acutus is an essential requisite to provide molecular reagents for functional genomic studies needed for elucidating mechanisms of action of toxins and surveying physiological events taking place in the very specialized secretory tissue. So this study provides a first global view of the genetic programs for the venom gland of Deinagkistrodon acutus described so far and an insight into molecular mechanism of toxin secreting.

Animals↗

Identification of human chromosome 22 transcribed sequences with ORF expressed sequence tags.

Transcribed sequences in the human genome can be identified with confidence only by alignment with sequences derived from cDNAs synthesized from naturally occurring mRNAs. We constructed a set of 250,000 cDNAs that represent partial expressed gene sequences and that are biased toward the central coding regions of the resulting transcripts. They are termed ORF expressed sequence tags (ORESTES). The 250,000 ORESTES were assembled into 81,429 contigs. Of these, 1, 181 (1.45%) were found to match sequences in chromosome 22 with at least one ORESTES contig for 162 (65.6%) of the 247 known genes, for 67 (44.6%) of the 150 related genes, and for 45 of the 148 (30.4%) EST-predicted genes on this chromosome. Using a set of stringent criteria to validate our sequences, we identified a further 219 previously unannotated transcribed sequences on chromosome 22. Of these, 171 were in fact also defined by EST or full length cDNA sequences available in GenBank but not utilized in the initial annotation of the first human chromosome sequence. Thus despite representing less than 15% of all expressed human sequences in the public databases at the time of the present analysis, ORESTES sequences defined 48 transcribed sequences on chromosome 22 not defined by other sequences. All of the transcribed sequences defined by ORESTES coincided with DNA regions predicted as encoding exons by genscan. (http://genes.mit.edu/GENSCAN.html).

Chromosomes, Human, Pair 22↗

Identification of novel highly expressed genes in pancreatic ductal adenocarcinomas through a bioinformatics analysis of expressed sequence tags.

In most microarray experiments, a significant fraction of the differentially expressed mRNAs identified correspond to expressed sequence tags (ESTs) and are generally discarded from further analyses. We used careful bioinformatics analyses to characterize those ESTs that were found to be highly overexpressed in a series of pancreatic adenocarcinomas. cDNA was prepared from 60 non-neoplastic samples (normal pancreas [n = 20], normal colon [n = 10], or normal duodenal mucosal [n = 30]) and from 64 pancreatic cancers (resected cancers [n = 50] or cancer cell lines [n = 14]) and hybridized to the complete Affymetrix Human Genome U133 GeneChip(R) set (arrays U133A and B) for simultaneous analysis of 45,000 fragments corresponding to 33,000 known genes and 6,000 ESTs. The GeneExpress(R) software system Fold Change Analysis Tool was used and 60 ESTs were identified that were expressed at levels at least 3-fold greater in the pancreatic cancers as compared to normal tissues. Searches against the human genomic sequence and comparative genomic analysis of human and mouse genomes was carried out using basic local alignment search tools (BLAST), BLASTN, and BLASTX, for identifying protein coding genes corresponding to the ESTs. Subsequently, in order to pick the most relevant candidate genes for a more detailed analysis, we looked for domains/motifs in the open reading frames using SMART and Pfam programs. We were able to definitively map 43 of the 60 ESTs to known or novel genes, and 15 of the ESTs could be localized in close proximity to a gene in the human genome although we were unable to establish that the EST was indeed derived from those genes. The differential expression of a subset of genes was confirmed at the protein level by immunohistochemical labeling of tissue microarrays (inhibin beta A [INHBA] and CD29) and/or at the transcript level by RT-PCR (INHBA, AKAP12, ELK3, FOXQ1, EIF5A2, and EFNA5). We conclude that bioinformatics tools can be used to characterize differentially overexpressed ESTs, and that some of these ESTs may represent diagnostically and therapeutically useful targets that might be missed using data solely from currently annotated databases.

Adenocarcinoma↗

Generating unigene collections of expressed sequence tag sequences for use in mass spectrometry identification.

Expressed sequence tag sequences remain the largest resource of DNA sequence for most organisms despite recent advances in genome sequencing. These sequences are short, fragmented versions of the expressed genes. By DNA sequence assembly, the fragments can be assembled into contiguous DNA sequences that are better suited for protein identification by mass spectrometry.

Cluster Analysis↗

Expressed sequence tags from eyestalk of kuruma prawn, Marsupenaeus japonicus.

We analyzed the expressed sequence tags (ESTs) obtained from a cDNA library of the eyestalk of the kuruma prawn, Marsupenaeus japonicus, to examine gene expression profile with special focus on female reproduction. The assembly of 1988 ESTs created 136 contigs from 738 ESTs; however 1250 ESTs remained singletons. Significant similarities (blast score > or = 50 bits) to the DNA sequences in the databank were found for only 16.7% of the 1386 sequences (136 contigs plus 1250 singletons), suggesting that the eyestalk library contains many unknown genes. Ribosomal RNA and mitochondrial respiration enzymes with significant similarities were found abundantly in the ESTs, whereas genes related to maturation or endocrine systems were scarce. Three ESTs were assumed to encode novel eyestalk hormones with marked similarities to pigment-dispersing hormone, molt-inhibiting hormone and crustacean hyperglycemic hormone. Sequences encoding a product highly homologous to farnesoic acid O-methyltransferase, an enzyme that produces methyl farnesoate, were also found.

Amino Acid Sequence↗

Gene discovery through expressed sequence Tag sequencing in Trypanosoma cruzi.

Analysis of expressed sequence tags (ESTs) constitutes a useful approach for gene identification that, in the case of human pathogens, might result in the identification of new targets for chemotherapy and vaccine development. As part of the Trypanosoma cruzi genome project, we have partially sequenced the 5' ends of 1, 949 clones to generate ESTs. The clones were randomly selected from a normalized CL Brener epimastigote cDNA library. A total of 14.6% of the clones were homologous to previously identified T. cruzi genes, while 18.4% had significant matches to genes from other organisms in the database. A total of 67% of the ESTs had no matches in the database, and thus, some of them might be T. cruzi-specific genes. Functional groups of those sequences with matches in the database were constructed according to their putative biological functions. The two largest categories were protein synthesis (23.3%) and cell surface molecules (10.8%). The information reported in this paper should be useful for researchers in the field to analyze genes and proteins of their own interest.

Animals↗

Comparative analysis of expressed sequence tags from cold-acclimated and non-acclimated leaves of Rhododendron catawbiense Michx.

An expressed sequence tag (EST) analysis approach was undertaken to identify major genes involved in cold acclimation of Rhododendron, a broad-leaf, woody evergreen species. Two cDNA libraries were constructed, one from winter-collected (cold-acclimated, CA; leaf freezing tolerance -53 degrees C) leaves, and the other from summer-collected (non-acclimated, NA; leaf freezing tolerance -7 degrees C) leaves of field-grown Rhododendron catawbiense plants. A total of 862 5'-end high-quality ESTs were generated by sequencing cDNA clones from the two libraries (423 from CA and 439 from NA library). Only about 6.3% of assembled unique transcripts were shared between the libraries, suggesting remarkable differences in gene expression between CA and NA leaves. Analysis of the relative frequency at which specific cDNAs were picked from each library indicated that four genes or gene families were highly abundant in the CA library including early light-induced proteins (ELIP), dehydrins/late embryogenesis abundant proteins (LEA), cytochrome P450, and beta-amylase. Similarly, seven genes or gene families were highly abundant in the NA library and included chlorophyll a/b-binding protein, NADH dehydrogenase subunit I, plastidic aldolase, and serine:glyoxylate aminotransferase, among others. Northern blot analyses for seven selected abundant genes confirmed their preferential expression in either CA or NA leaf tissues. Our results suggest that osmotic regulation, desiccation tolerance, photoinhibition tolerance, and photosynthesis adjustment are some of the key components of cold adaptation in Rhododendron.

Acclimatization↗

[Identification and analysis of expressed sequence tags related to K562 cells into erythroid differentiation].

OBJECTIVE: To isolate expressed sequence tags (ESTs) related to K562 cells erythroid differentiation. METHODS: Modified differential display reverse transcription polymerase chain reaction (DDRT-PCR) method was applied to identify differential ESTs in uninduced and induced K562 cells by HEMIN for 36 hours. Remarkable differential ESTs were firstly selected for cloning, sequencing and bioinformational analyzing. Several ESTs representing new sequence or providing functional clue were selected for Northern blot analysis. RESULTS: Sixty differentially expressed cDNA fragments related to K562 cells inducted into erythroid differentiation by HEMIN were obtained. Among them, 38 were upregulated and 22 downregulated. Among the 40 differential ESTs selected for cloning, sequencing and bioinformationally analyzing, 23 were found to match to known GenBank sequences and 10 represented cDNA sequences with only dbEST database matches and 7 ESTs have no any database matches. The results of 6 in 8 ESTs selected for Northern blot analysis were shown to be consistent with the differential expressions of DDRT-PCR. CONCLUSIONS: The improved DDRT-PCR method had successfully overcome the problem of false positive. These ESTs provide some clue for studying the molecular mechanisms and regulation network of erythroid differentiation.

Cell Differentiation↗

Analysis of expressed sequence tags from oil palm (Elaeis guineensis).

This is the first report of a systematic study of genes expressed by means of expressed sequence tag (EST) analysis in oil palm, a species of the Arecales order, a phylogenetically key clade of monocotyledons that is not widely represented in the sequence databases. Five different cDNA libraries were generated from male and female inflorescences, shoot apices and zygotic embryos and unidirectional systematic sequencing was performed. A total of 2411 valid EST sequences were thus obtained. Cluster analysis enabled the identification of 209 groups of related sequences and 1874 singletons. Putative functions were assigned to 1252 of the set of 2083 non-redundant ESTs obtained. The EST database described here is a first step towards gene discovery and cDNA array-based expression analysis in oil palm.

Arecaceae↗

Expressed sequence tags from the plant trypanosomatid Phytomonas serpens.

We have generated 2190 expressed sequence tags (ESTs) from a cDNA library of the plant trypanosomatid Phytomonas serpens. Upon processing and clustering the set of 1893 accepted sequences was reduced to 697 clusters consisting of 452 singletons and 245 contigs. Functional categories were assigned based on BLAST searches against a database of the eukaryotic orthologous groups of proteins (KOG). Thirty six percent of the generated sequences showed no hits against the KOG database and 39.6% presented similarity to the KOG classes corresponding to translation, ribosomal structure and biogenesis. The most populated cluster contained 45 ESTs homologous to members of the glucose transporter family. This fact can be immediately correlated to the reported Phytomonas dependence on anaerobic glycolytic ATP production due to the lack of cytochrome-mediated respiratory chain. In this context, not only a number of enzymes of the glycolytic pathway were identified but also of the Krebs cycle as well as specific components of the respiratory chain. The data here reported, including a few hundred unique sequences and the description of tandemly repeated motifs and putative transcript stability motifs at untranslated mRNA ends, represent an initial approach to overcome the lack of information on the molecular biology of this organism.

Amino Acid Sequence↗

Expressed sequence tag (EST) analysis of a Schistosoma japonicum cercariae cDNA library.

Expressed sequence tags (ESTs) constitute a rapid and informative strategy for studying gene-expression profiles of specific stages of schistosomes. To date, only approximately equal 2000 ESTs of Schistosoma japonicum have been deposited in databases. This is insufficient to understand the biology and development of this species. In this report, a cDNA library constructed from S. japonicum cercariae RNA was used to generate ESTs. Cercariae are the larval forms of Schistosoma responsible for infection of the vertebrate host and one of the main objectives of this research was to discover and characterize new and unique genes from this stage. The expression products of those stage-specific genes can potentially be useful as new drugs or vaccine targets applicable for controlling Asian schistosomiasis. In our study, 101 cDNA clones were sequenced either from 5' or 3' end of the cDNAs. Some 42 ESTs (42%) matched known genes, while 59 ESTs did not match with any known genes. Among the 42 former ESTs, 29 (informative ESTs) matched to functional genes and 13 matched with ribosomal proteins or RNA genes. Among the latter 59 ESTs, 21 matched with published ESTs of S. japonicum or S. mansoni, two matched with human ESTs and the other 18 did not match with any published sequences. The informative ESTs could be grouped into nine categories: regulatory and signaling proteins (24.1%), transcription and translation machinery proteins (13.8%), RNA binding proteins (6.9%), structural and cytoskeletal proteins (6.9%), DNA binding proteins (3.4%), DNA scaffold proteins (3.4%), transporter proteins (3.4%) and others (31%). Some functional genes relevant to the physiology of cercariae are discussed.

Animals↗

2058 expressed sequence tags (ESTs) from a human fetal lung cDNA library.

ESTs (expressed sequence tags) provide complementary resources for structural and functional analyses of the human genome. We have performed single-pass sequencing of 2058 randomly selected, directionally cloned cDNAs isolated from a fetal-lung cDNA library constructed with oligo(dT) primers. Computer analyses of the 5'-end sequences revealed that 60.4% of the clones were considered to be identical to previously reported human genes or ESTs; 9.0% of them showed significant homology to known genes in human, other mammals, or lower organisms; 30.6% showed no homology to any genes or DNA sequences in the public database. These data and reagents will be useful for future investigations of gene expression during prenatal development of human lung.

Amino Acid Sequence↗

Gene expression in the salivary complexes from Haementeria depressa leech through the generation of expressed sequence tags.

A survey of the transcriptional profile of Haementeria depressa Ringuelet, 1972 (Annelida, Hirudinea) salivary complexes was produced through expressed sequence tag (EST). Sequences from 898 independent clones were assembled in 555 clusters, representing the transcript profile of this tissue. The repertoire of possible proteins involved in feeding and host interaction processes of the leech corresponded to 10.6% of all identified transcripts (67 clusters), being the carbonic anhydrases (30%), several coagulation inhibitors (25%) and hemerythrin-like molecules (19%), the major components. Among the 387 clusters matching cellular proteins, the majority represents molecules involved in gene and protein expression, reflecting a high specialization of this tissue for protein synthesis. Our H. depressa dbEST was also compared to those from other blood-feeding organisms, providing evidences that among the secreted proteins, the coagulation inhibitors present a profile very characteristic of this animal class.

Amino Acid Sequence↗

Generation of a total of 6483 expressed sequence tags from 60 day-old bovine whole fetus and fetal placenta.

Expressed sequence tags (ESTs) generated based on characterization of clones isolated randomly from cDNA libraries are used to study gene expression profiles in specific tissues and to provide useful information for characterizing tissue physiology. In this study, two directionally cloned cDNA libraries were constructed from 60 day-old bovine whole fetus and fetal placenta. We have characterized 5357 and 1126 clones, and then identified 3464 and 795 unique sequences for the fetus and placenta cDNA libraries: 1851 and 504 showed homology to already identified genes, and 1613 and 291 showed no significant matches to any of the sequences in DNA databases, respectively. Further, we found 94 unique sequences overlapping in both the fetus and the placenta, leading to a catalog of 4165 genes expressed in 60 day-old fetus and placenta. The catalog is used to examine expression profile of genes in 60 day-old bovine fetus and placenta.

Animals↗

Mining for single nucleotide polymorphisms and insertions/deletions in maize expressed sequence tag data.

We have developed a computer based method to identify candidate single nucleotide polymorphisms (SNPs) and small insertions/deletions from expressed sequence tag data. Using a redundancy-based approach, valid SNPs are distinguished from erroneous sequence by their representation multiple times in an alignment of sequence reads. A second measure of validity was also calculated based on the cosegregation of the SNP pattern between multiple SNP loci in an alignment. The utility of this method was demonstrated by applying it to 102,551 maize (Zea mays) expressed sequence tag sequences. A total of 14,832 candidate polymorphisms were identified with an SNP redundancy score of two or greater. Segregation of these SNPs with haplotype indicates that candidate SNPs with high redundancy and cosegregation confidence scores are likely to represent true SNPs. This was confirmed by validation of 264 candidate SNPs from 27 loci, with a range of redundancy and cosegregation scores, in four inbred maize lines. The SNP transition/transversion ratio and insertion/deletion size frequencies correspond to those observed by direct sequencing methods of SNP discovery and suggest that the majority of predicted SNPs and insertion/deletions identified using this approach represent true genetic variation in maize.

Base Sequence↗

An expressed sequence tag analysis of the life-cycle of the parasitic nematode Strongyloides ratti.

14,761 expressed sequence tags (ESTs) were generated, representing five stages during the parasitic and free-living phases of the life-cycle of the parasitic nematode Strongyloides ratti. These ESTs formed 4152 clusters, of which 97% contained 10 or fewer ESTs and 66% were singletons. These 4152 clusters are likely to represent approximately 20% of S. ratti's genes. The clusters' consensus sequences were used to assign each cluster to one of three databases: (i) Caenorhabditis elegans and C. briggsae sequences; (ii) other nematode sequences; (iii) non-nematode sequences. This approach has identified putative nematode-specific genes, that may be targets for developing approaches for parasitic nematode control. Approximately 25% of the clusters have no significant alignments and may therefore represent novel genes. The EST representation between the libraries was used to analyse stage-specific or -biased expression in silico. This showed that 81% of clusters are present in only one library and 12% are present in any two libraries, indicating substantial stage-specificity of gene expression. The 30-most abundantly expressed clusters were analysed in further detail. Many of these have significantly different parasitic- or free-living-specific or -biased expression. Many of the parasitic-specific genes are, as yet, uncharacterised: one of these represents 25% of all ESTs obtained from the parasitic stage.

Animals↗