PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Expressed Sequence Tags”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Identification of candidate coding region single nucleotide polymorphisms in 165 human genes using assembled expressed sequence tags.

Using assembled expressed sequence tags (ESTs) from 50 different cDNA libraries, we have identified contigs that represent the complete coding sequences of 850 known human genes, and have scanned these for high quality sequence substitutions. We report the identification and characteristics of 201 candidate single nucleotide polymorphisms found in the coding sequences (cSNPs) of 165 of these genes. Using a conservative calculation, coding region nucleotide diversity (the average number of differences between any pair of chromosomes) was found to be 3 per 10,000 bp based on this data. This analysis reveals that assembled ESTs from multiple libraries may provide a rich source of comparative sequences to search for cSNPs in the human genome.

Amino Acid Substitution↗

Linkage mapping and comparative analysis of bovine expressed sequence tags (ESTs).

Bovine expressed sequence tags (ESTs) containing microsatellites are suitable markers for both linkage and comparative maps. We isolated clones from a bovine fetal thigh skeletal muscle cDNA library that were positive for a (CA)10 probe. Thirty individual clones were isolated and characterised by sequencing. Sequences from the 5' and 3' ends of a clone were considered as separate ESTs until a contiguous sequence was identified. A total of 47 ESTs were sequenced from the 5' and/or 3' ends and full sequence was obtained for the 30 clones. BLAST nucleotide analysis identified significant homology to known mammalian coding regions for 31 of the bovine ESTs, 30 of which also matched human ESTs or sequence-tagged sites (STS). The remaining 16 bovine ESTs represented novel transcripts. Microsatellites were isolated in 27 of the ESTs, 11 of which were developed into markers and placed on the MARC bovine linkage map. Human cytogenetic map positions were available for 20 of the 30 human EST orthologs, and a putative bovine map position for 17 of the sequences could be inferred using comparative mapping data. These results demonstrated that mapping bovine ESTs containing microsatellites is a plausible strategy to increase the density of gene markers on the bovine linkage and comparative maps.

Animals↗

A hitchhiker's guide to expressed sequence tag (EST) analysis.

Expressed sequence tag (EST) sequencing projects are underway for numerous organisms, generating millions of short, single-pass nucleotide sequence reads, accumulating in EST databases. Extensive computational strategies have been developed to organize and analyse both small- and large-scale EST data for gene discovery, transcript and single nucleotide polymorphism analysis as well as functional annotation of putative gene products. We provide an overview of the significance of ESTs in the genomic era, their properties and the applications of ESTs. Methods adopted for each step of EST analysis by various research groups have been compared. Challenges that lie ahead in organizing and analysing the ever increasing EST data have also been identified. The most appropriate software tools for EST pre-processing, clustering and assembly, database matching and functional annotation have been compiled (available online from http://biolinfo.org/EST). We propose a road map for EST analysis to accelerate the effective analyses of EST data sets. An investigation of EST analysis platforms reveals that they all terminate prior to downstream functional annotation including gene ontologies, motif/pattern analysis and pathway mapping.

Animals↗

Identification of differentially expressed genes in normal and malignant prostate by electronic profiling of expressed sequence tags.

Differentially expressed genes between corresponding normal and cancertissue can advance our understanding of the molecular basis of malignancy and potentially serve as biomarkers or prognostic markers of malignancy. To identify differentially expressed genes in prostate cancer, we used a procedure combining electronic expression profiling of the prostate expressed sequence tag (EST) database and molecular biology techniques. A novel electronic expression-profiling algorithm was developed to search publicly available EST sequences for genes that show significant differential expression in prostate cancer compared with normal prostate tissue. Approximately 600 genes expressed in prostate were identified through adequate EST counts of ESTs for electronic profiling. Of these 600 genes, 9 showed statistically significant differences in their EST counts between cancer and normal prostate and were further analyzed. The predictions associated with electronic profiling were experimentally verified for two genes, cysteine-rich secretory protein 3 (CRISP-3) and deadenylating nuclease (DAN), using real-time reverse transcription-PCR with total RNA extracted from cells isolated by laser capture microdissection. In five of five Gleason score 6 cancer cases, CRISP-3 expression was increased >50 fold, whereas the expression of DAN was reduced by >80%.

Algorithms↗

Odorant receptor expressed sequence tags demonstrate olfactory expression of over 400 genes, extensive alternate splicing and unequal expression levels.

BACKGROUND: The olfactory receptor gene family is one of the largest in the mammalian genome. Previous computational analyses have identified approximately 1,500 mouse olfactory receptors, but experimental evidence confirming olfactory function is available for very few olfactory receptors. We therefore screened a mouse olfactory epithelium cDNA library to obtain olfactory receptor expressed sequence tags, providing evidence of olfactory function for many additional olfactory receptors, as well as identifying gene structure and putative promoter regions. RESULTS: We identified more than 1,200 odorant receptor cDNAs representing more than 400 genes. Using real-time PCR to confirm expression level differences suggested by our screen, we find that transcript levels in the olfactory epithelium can differ between olfactory receptors by up to 300-fold. Differences for one gene pair are apparently due to both unequal numbers of expressing cells and unequal transcript levels per expressing cell. At least two-thirds of olfactory receptors exhibit multiple transcriptional variants, with alternative isoforms of both 5' and 3' untranslated regions. Some transcripts (5%) utilize splice sites within the coding region, contrary to the stereotyped olfactory receptor gene structure. Most atypical transcripts encode nonfunctional olfactory receptors, but can occasionally increase receptor diversity. CONCLUSIONS: Our cDNA collection confirms olfactory function of over one-third of the intact mouse olfactory receptors. Most of these genes were previously annotated as olfactory receptors based solely on sequence similarity. Our finding that different olfactory receptors have different expression levels is intriguing given the one-neuron, one-gene expression regime of olfactory receptors. We provide 5' untranslated region sequences and candidate promoter regions for more than 300 olfactory receptors, valuable resources for computational regulatory motif searches and for designing olfactory receptor microarrays and other experimental probes.

Alternative Splicing↗

A new dynamic tool to perform assembly of expressed sequence tags (ESTs).

MOTIVATION: Expressed Sequence Tags (ESTs) are short single-pass DNA sequences obtained from either ends of cDNA clones. To exploit these sequences efficiently, a dynamic Web-tool has been developed which uses these data to perform fast virtual cloning of cDNAs. RESULTS: Starting with a query sequence, the user is able to identify related ESTs and extend the sequence of interest step by step, possibly to a full-length transcript. Graphical views of the clustering are used to monitor the progress of a particular 'cloning' project. Potential open reading frames are detected by positional base preference, and hyperlinks to other Worldwide Web sites allows the user to retrieve information relevant to each EST in a cluster (e.g. sequence traces, clone size, plate position). Apart from cDNA cloning, this tool also provides a mechanism for collating gene families and polymorphism sites.

Algorithms↗

Application of representational difference analysis to identify sequence tags expressed by Metarhizium anisopliae during the infection process of the tick Boophilus microplus cuticle.

Metarhizium anisopliae is a well-characterized biocontrol agent of a wide range of plagues, including insects and acari. To identify genes involved in the infection process, representational difference analysis was performed using cDNA generated from germinated conidia of M. anisopliae in the tick Boophilus microplus cuticle, and cDNA generated during fungal growth in glucose-rich medium. Sequence determination of approximately 135 clones and comparison analysis using public databases led to the identification of 34 sequences and 14 expressed sequence tags with known orthologs. As expected, almost all identified sequences showed significant similarity to other fungal genes. The diversity of gene clusters found reflects the participation of several proteins in the early infection process of M. anisopliae in the cattle tick B. microplus.

Animals↗

Gene discovery and gene expression in the rice blast fungus, Magnaporthe grisea: analysis of expressed sequence tags.

Over 28,000 expressed sequence tags (ESTs) were produced from cDNA libraries representing a variety of growth conditions and cell types. Several Magnaporthe grisea strains were used to produce the libraries, including a nonpathogenic strain bearing a mutation in the PMK1 mitogen-activated protein kinase. Approximately 23,000 of the ESTs could be clustered into 3,050 contigs, leaving 5,127 singleton sequences. The estimate of 8,177 unique sequences indicates that over half of the genes of the fungus are represented in the ESTs. Analysis of EST frequency reveals growth and cell type-specific patterns of gene expression. This analysis establishes criteria for identification of fungal genes involved in pathogenesis. A large fraction of the genes represented by ESTs have no known function or described homologs. Manual annotation of the most abundant cDNAs with no known homologs allowed us to identify a family of metallothionein proteins present in M. grisea, Neurospora crassa, and Fusarium graminearum. In addition, multiply represented ESTs permitted the identification of alternatively spliced mRNA species. Alternative splicing was rare, and in most cases, the alternate mRNA forms were unspliced, although alternative 5' splice sites were also observed.

Expressed Sequence Tags↗

Identification of differentially expressed transcripts from maturing stem of sugarcane by in silico analysis of stem expressed sequence tags and gene expression profiling.

Sugarcane accumulates high concentrations of sucrose in the mature stem and a number of physiological processes on-going in maturing stem tissue both directly and indirectly allow this process. To identify transcripts that are associated with stem maturation, we compared patterns of gene expression in maturing and immature stem tissue by expression profiling and bioinformatic analysis of sets of stem ESTs. This study complements a previous study of gene expression associated directly with sugar metabolism in sugarcane. A survey of sequences derived from stem tissue identified an abundance of several classes of sequence that are associated with fibre biosynthesis in the maturing stem. A combination of EST analyses and microarray hybridization revealed that genes encoding homologues of the dirigent protein, a protein that assists in the stereospecificity of lignin assembly, were the most abundant and most strongly differentially expressed transcripts in maturing stem tissue. There was also evidence of coordinated expression of other categories of fibre biosynthesis and putative defence- and stress-related transcripts in the maturing stem. This study has demonstrated the utility of genomic approaches using large-scale EST acquisition and microarray hybridization techniques to highlight the very significant transcriptional investment the maturing stem of sugarcane has placed in fibre biosynthesis and stress tolerance, in addition to its already well-documented role in sugar accumulation.

Amino Acid Sequence↗

Ciona intestinalis cDNA projects: expressed sequence tag analyses and gene expression profiles during embryogenesis.

Ascidians are primitive chordates. Their fertilized egg develops quickly into a tadpole-type larva, which consists of a small number but distinct types of cells, including those of epidermis, central nervous system with two sensory organs, endoderm and mesenchyme in the trunk, and notochord and muscle in the tail. This configuration of the ascidian tadpole is thought to represent the most simplified and primitive chordate body plan. In addition, the free-swimming and non-feeding larvae metamorphose into sessile and filter-feeding adults. The genome size of Ciona intestinalis is estimated to be about 160 Mb, and the number of genes approximately 15,500. The present Ciona cDNA projects focused on gene expression profiles of fertilized eggs, 32-110-cell stage embryos, tailbud embryos, larvae, and young adults. Expressed sequence tags (ESTs) of the 5'-most end and 3'-most end of more than 3000 clones were determined at each developmental stage, and the clones were categorized into independent clusters using the 3'-end sequences. Nearly 1000 clusters of them were then analyzed in detail of their sequences against a BLASTX search. This analysis demonstrates that, on average, half of the clusters showed proteins with sequence similarities to known proteins and the other half did not show sequence similarities to known proteins. Genes with sequence similarities were further categorized into three major subclasses, depending on their functions. Furthermore, the expression profiles of all of the clusters were analyzed by whole-mount in situ hybridization. This analysis highlights gene expression patterns characteristic to each developmental stage. As a result, the present study provides many new molecular markers for each of the tissues and/or organs that constitutes the Ciona tailbud embryo. This sequence information will be used for further comparative genome studies to explore molecular mechanisms involved in the formation of one of the most primitive chordate body plans. All of the data fully characterized may be viewed at the web site http://ghost.zool.kyoto-u.ac.jp.

Animals↗

Identification of beta-amyrin and sophoradiol 24-hydroxylase by expressed sequence tag mining and functional expression assay.

Triterpenes exhibit a wide range of structural diversity produced by a sequence of biosynthetic reactions. Cyclization of oxidosqualene is the initial origin of structural diversity of skeletons in their biosynthesis, and subsequent regio- and stereospecific hydroxylation of the triterpene skeleton produces further structural diversity. The enzymes responsible for this hydroxylation were thought to be cytochrome P450-dependent monooxygenase, although their cloning has not been reported. To mine these hydroxylases from cytochrome P450 genes, five genes (CYP71D8, CYP82A2, CYP82A3, CYP82A4 and CYP93E1) reported to be elicitor-inducible genes in Glycine max expressed sequence tags (EST), were amplified by PCR, and screened for their ability to hydroxylate triterpenes (beta-amyrin or sophoradiol) by heterologous expression in the yeast Saccharomyces cerevisiae. Among them, CYP93E1 transformant showed hydroxylating activity on both substrates. The products were identified as olean-12-ene-3beta,24-diol and soyasapogenol B, respectively, by GC-MS. Co-expression of CYP93E1 and beta-amyrin synthase in S. cerevisiae yielded olean-12-ene-3beta,24-diol. This is the first identification of triterpene hydroxylase cDNA from any plant species. Successful identification of a beta-amyrin and sophoradiol 24-hydroxylase from the inducible family of cytochrome P450 genes suggests that other triterpene hydroxylases belong to this family. In addition, substrate specificity with the obtained P450 hydroxylase indicates the two possible biosynthetic routes from triterpene-monool to triterpene-triol.

Base Sequence↗

High-throughput GLGI procedure for converting a large number of serial analysis of gene expression tag sequences into 3' complementary DNAs.

Serial analysis of gene expression (SAGE) is a powerful technique for genome-wide analysis of gene expression. However, two-thirds of SAGE tags cannot be used directly for gene identification for two reasons. First, many SAGE tags match several known expressed sequences, owing to the short length of SAGE tag sequences. Second, many SAGE tags do not match any known expressed sequences, presumably because the sequences corresponding to these SAGE tags have not been identified. These two problems can be solved by extension of the SAGE tags into 3' complementary DNAs (cDNAs) by use of the GLGI technique (generation of longer cDNA fragments from SAGE tags for gene identification). We have improved the original GLGI technique into a high-throughput procedure for simultaneous conversion of a large number of SAGE tags into corresponding 3' cDNAs. The whole process is simple, rapid, low-cost, and highly efficient, as shown by our use of this procedure for analyzing hundreds of SAGE tags. In addition to identifying the correct gene for SAGE tags with multiple matches, GLGI can be used for large-scale identification of novel genes by converting novel SAGE tags into 3' cDNAs. Applying this high-throughput procedure should accelerate the rate of gene identification significantly in the human and other eukaryotic genomes.

Animals↗

Expressed sequence tags: analysis and annotation.

Expressed sequence tags (ESTs) present a special set of problems for bioinformatic analysis. They are partial and error-prone, and large datasets can have significant internal redundancy. To facilitate analysis of small EST datasets from in-house projects, we present an integrated "pipeline" of tools that take EST data from sequence trace to database submission. These tools also can be used to provide clustering of ESTs into putative genes and to annotate these genes with preliminary sequence similarity searches. The systems are written to use the public-domain LINUX environment and other openly available analytical tools.

Computational Biology↗

An expressed sequence tag survey of gene expression in the pond snail Lymnaea stagnalis, an intermediate vector of trematodes [corrected].

The pond snail Lymnaea stagnalis is an intermediate vector for the liver fluke Fasciola hepatica, a common parasite of ruminants and humans. Yet, despite being a disease of medical and economic importance, as well as a potentially useful comparative tool, the genetics of the relationship between Lymnaea and Fasciola has barely been investigated. As a complement to forthcoming F. hepatica expressed sequence tags (ESTs), we generated 1320 ESTs from L. stagnalis central nervous system (CNS) libraries. We estimate that these sequences derive from 771 different genes, of which 374 showed significant similarity to proteins in public databases, and 169 were similar to ESTs from the snail vector Biomphalaria glabrata. These L. stagnalis ESTs will provide insight into the function of the snail CNS, as well as the molecular components of behaviour and response to parasitism. In the future, the comparative analysis of Lymnaea/Fasciola with Biomphalaria/Schistosoma will help to understand both conserved and divergent aspects of the host-parasite relationship. The L. stagnalis ESTs will also assist gene prediction in the forthcoming B. glabrata genome sequence. The dataset is available for searching on the world-wide web at http://zeldia.cap.ed.ac.uk/mollusca.html.

Amino Acid Sequence↗

Comparative expressed-sequence-tag analysis of differential gene expression profiles in PC-12 cells before and after nerve growth factor treatment.

Nerve growth factor-induced differentiation of adrenal chromaffin PC-12 cells to a neuronal phenotype involves alterations in gene expression and represents a model system to study neuronal differentiation. We have used the expressed-sequence-tag approach to identify approximately 600 differentially expressed mRNAs in untreated and nerve growth factor-treated PC-12 cells that encode proteins with diverse structural and biochemical functions. Many of these mRNAs encode proteins belonging to cellular pathways not previously known to be regulated by nerve growth factor. Comparative expressed-sequence-tag analysis provides a basis for surveying global changes in gene-expression patterns in response to biological signals at an unprecedented scale, is a powerful tool for identifying potential interactions between different cellular pathways, and allows the gene-expression profiles of individual genes belonging to a particular pathway to be followed.

Animals↗

The poplar root transcriptome: analysis of 7000 expressed sequence tags.

To date, most poplar expressed sequence tags (ESTs) are from above-ground tissues such as wood, leaf and buds. Here, we present a large-scale production of ESTs from roots of the hybrid cottonwood, Populus trichocarpaxdeltoides. cDNA libraries were generated from the root system of 2-month-old rooted cuttings, and roots of 2.5-month-old cuttings water-stressed for 19 days. Partial sequences obtained from 7013 clones were assembled into 1347 clusters and 3527 singletons. This set of ESTs represents 4874 unique transcripts expressed in roots. Putative functions could be assigned to 3021 (62%) of the transcripts. A significant portion of the ESTs encode proteins of common metabolic pathways; energy and metabolism represented 5% and 8% of total transcripts, respectively. Of specific interest to root functions are the 6% of ESTs involved in signalling pathways and hormone metabolism, and 4% encoding transporters and channels. The current poplar root ESTs and the aspen root ESTs present in public databases represent 6700 unique transcripts. The Unigene set was selected from the ESTs and used to generate nylon microarrays. Changes in aquaporins and transporter transcripts were then studied during adventitious root development.

Aquaporins↗

Expression profile of two storage-protein gene families in hexaploid wheat revealed by large-scale analysis of expressed sequence tags.

To discern expression patterns of individual storage-protein genes in hexaploid wheat (Triticum aestivum cv Chinese Spring), we analyzed comprehensive expressed sequence tags (ESTs) of common wheat using a bioinformatics technique. The gene families for alpha/beta-gliadins and low molecular-weight glutenin subunit were selected from the EST database. The alignment of these genes enabled us to trace the single nucleotide polymorphism sites among both genes. The combinations of single nucleotide polymorphisms allowed us to assign haplotypes into their homoeologous chromosomes by allele-specific PCR. Phylogenetic analysis of these genes showed that both storage-protein gene families rapidly diverged after differentiation of the three genomes (A, B, and D). Expression patterns of these genes were estimated based on the frequencies of ESTs. The storage-protein genes were expressed only during seed development stages. The alpha/beta-gliadin genes exhibited two distinct expression patterns during the course of seed maturation: early expression and late expression. Although the early expression genes among the alpha/beta-gliadin and low molecular-weight glutenin subunit genes showed similar expression patterns, and both genes from the D genome were preferentially expressed rather than those from the A or B genome, substantial expression of two early expression genes from the A genome was observed. The phylogenetic relationships of the genes and their expression patterns were not correlated. These lines of evidence suggest that expression of the two storage-protein genes is independently regulated, and that the alpha/beta-gliadin genes possess novel regulation systems in addition to the prolamin box.

Base Sequence↗

Toxoplasma gondii expressed sequence tags: insight into tachyzoite gene expression.

Analysis of DNA sequences from the 5' end of 239 directionally cloned Toxoplasma gondii RH strain tachyzoite-derived cDNAs revealed significant similarity to several classes of genes/proteins including 24 ribosomal proteins, five metabolic enzymes, four cell-cycle regulators and 15 previously cloned T. gondii genes. The remaining sequences with no significant match include several which were recovered more than once. The variety and redundancy of expressed sequence tags (ESTs GenBank accession numbers T62239-T62475) in this sample suggest that the tachyzoite cDNA library reflects tachyzoite gene expression. A large scale EST effort should uncover many new genes and provide a wealth of information about genes involved with the growth and proliferation of tachyzoites.

Animals↗