PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Computational Proteomics Analysis System (CPAS): an extensible, open-source analytic system for evaluating and publishing proteomic data and high throughput biological experiments.

The open-source Computational Proteomics Analysis System (CPAS) contains an entire data analysis and management pipeline for Liquid Chromatography Tandem Mass Spectrometry (LC-MS/MS) proteomics, including experiment annotation, protein database searching and sequence management, and mining LC-MS/MS peptide and protein identifications. CPAS architecture and features, such as a general experiment annotation component, installation software, and data security management, make it useful for collaborative projects across geographical locations and for proteomics laboratories without substantial computational support.

Computational Biology↗

The FSSP database of structurally aligned protein fold families.

FSSP (families of structurally similar proteins) is a database of structural alignments of proteins in the Protein Data Bank (PDB). The database currently contains an extended structural family for each of 330 representative protein chains. Each data set contains structural alignments of one search structure with all other structurally significantly similar proteins in the representative set (remote homologs, < 30% sequence identity), as well as all structures in the Protein Data Bank with 70-30% sequence identity relative to the search structure (medium homologs). Very close homologs (above 70% sequence identity) are excluded as they rarely have marked structural differences. The alignments of remote homologs are the result of pairwise all-against-all structural comparisons in the set of 330 representative protein chains. All such comparisons are based purely on the 3D co-ordinates of the proteins and are derived by automatic (objective) structure comparison programs. The significance of structural similarity is estimated based on statistical criteria. The FSSP database is available electronically from the EMBL file server and by anonymous ftp (file transfer protocol).

Amino Acid Sequence↗

Analysis of the mouse proteome. (I) Brain proteins: separation by two-dimensional electrophoresis and identification by mass spectrometry and genetic variation.

The total protein of the mouse brain was fractionated into three fractions, supernatant, pellet extract and rest pellet suspension, by a procedure that avoids any loss of groups or classes of proteins. The supernatant proteins were resolved to a maximum by large-gel two-dimensional electrophoresis. Two-dimensional patterns from ten individual mice of the commonly used inbred strain C57BL/6 (species: Mus musculus) were prepared. The master pattern was subjected to densitometry, computer-assisted image analysis and treatment with our spot detection program. The resulting two-dimensional pattern, a standard pattern for mouse brain supernatant proteins, was divided into 40 squares, calibrated, and specified by providing each spot with a number. The complete pattern and each of the 40 squares are shown in our homepage (http://www.charite.de/ humangenetik). The standard pattern comprises 8767 protein spots. To identify the proteins known so far in the brain fraction investigated, a first set of 200 spots was analyzed by matrix-assisted laser desorption/ionization - mass spectrometry (MALDI-MS) after in-gel digestion. By screening protein databases 115 spots were identified; by extending the analysis to selected, genetically variant protein spots, 166 spots (including some spot series) were identified in total. This number was increased to 331 by adding protein spots identified indirectly by a genetic approach. By comparing the two-dimensional patterns from C57BL/6 mice with those of another mouse species (Mus spretus), more than 1000 genetically variant spots were detected. The genetic analysis allowed us to recognize spot families, i.e., protein spots that represent the same protein but that are post-translationally modified. If some members of the family were identified, the whole family was considered as being identified. Spot families were investigated in more detail, and interpreted as the result of protein modification or degradation. Genetic analysis led to the interesting finding that the size of spot families, i.e., the extent of modification or degradation of a protein, can be genetically determined. The investigation presented is a first step towards a systematic analysis of the proteome of the mouse. Proteome analysis was shown to become more efficient, and, at the same time, linked to the genome, by combining protein analytical and genetic methods.

Animals↗

Micropreparative immobilized pH gradient two-dimensional electrophoresis in combination with protein microsequencing for the analysis of human liver proteins.

Simplified methodology has been developed for the direct N-terminal amino acid microsequencing of human liver and hepatoma derived polypeptides, following micropreparative two-dimensional polyacrylamide gel electrophoresis (2-D PAGE). Utilization of immobilized pH gradient (IPG) gel strips in the first dimension permitted protein loading of 0.5-2.0 mg with negligible diminution of polypeptide resolution. Following 2-D separation and electrotransfer to polyvinylidene difluoride (PVDF) membranes nearly 100 well resolved Ponceau S stained polypeptides were readily visualized, from which, 32 adult liver S-9 and 72 HepG2 nuclear cytosolic polypeptides were subjected to N-terminal microsequencing. Twenty normal adult liver and 54 HepG2 polypeptides yielded N-terminal sequence information, of which 17 and 19 polypeptides, respectively, exhibited high sequence homology to previously identified proteins. The initial yields of the proteins sequenced ranged from 2-14 pmols and yielded sequences of 14-26 amino acid residues. Many of the adult liver and HepG2 proteins contained inferred leader sequences since the first sequenced residue was several (20-30) residues from the methionine initiation site predicted by the cDNA of the adult liver. Quantitative comparison of 60 well characterized hepatic proteins between normal adult liver and two nontransformed, Chang and WRL-68, and four human hepatoma derived cell lines, HepG2, Huh-7, FOCUS, and SK-Hep, revealed a high homogeneity of protein expression both qualitatively and quantitatively in both whole cell lysate and purified nuclear preparations. Most notable differences include the previously characterized polypeptides: carbamoyl phosphate synthase, MER5 homologous protein, cytidylate kinase, phosphatidylethanolamine-binding protein and mitochondrial enoyl-CoA hydratase as well as three N-terminally blocked polypeptides: 11 (63 kDa/pI 7.00), 56 (26/6.45) and 59 (22/6.00) all of which were expressed at similar levels in normal adult liver tissue and each of the nontransformed, Chang and WRL-68, cell lines but not expressed or expressed at greatly decreased levels in each of tumor derived liver cell lines. Pyruvate carboxylase, superoxide dismutase, serotransferrin, liver fatty acid binding protein, 1-hydroxyprostaglandin dehydrogenase, NADH dehydrogenase (ubiquinone) as well as three N-terminally blocked polypeptides: 9 (57/6.00), 53 (24/4.90) and 63 (16/4.70) were detected only in whole adult liver tissue and not in any of the cultured cell lines. Two additional polypeptides: U35, (27/6.05) and 58 (22/5.70) yielded N-terminal partial amino acid sequences but were not identified in established protein databases. We have shown that micropreparative IPG 2-D PAGE In combination with protein microsequencing provides a convenient one step procedure to rapidly obtain partial amino acid sequence information for nearly 100 individual polypeptides directly from a single 2-D PAGE gel with numerous applications to a wide variety of biological model systems.

Adult↗

Profiling the progression of cancer: separation of microsomal proteins in MCF10 breast epithelial cell lines using nonporous chromatophoresis.

The heterogeneity of cellular protein expression has stimulated development of separations targeting smaller groups of related proteins rather than entire proteomes. The following work describes the development of a technique for the characterization of membrane subproteomes from five different breast epithelial cell lines. Intact membrane proteins are separated by hydrophobicity in the first dimension using nonporous reversed-phase high-performance liquid chromatography (RP-HPLC) to generate unique chromatographic profiles. Fractions of eluent are further separated using sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) to create distinct banding patterns. This hybrid liquid phase/gel phase method circumvents issues of membrane protein precipitation and provides a simple strategy aimed at isolating and characterizing a traditionally underrepresented protein class. Membrane protein profiles are created that discriminate between microsomal fractions of breast epithelial cells in different stages of neoplastic progression. Proteins are subsequently identified using matrix-assisted laser desorption/ionization - mass spectrometry (MALDI-MS) mass fingerprinting and MALDI-quadrupole time of flight - tandem mass spectrometry (QTOF-MS/MS) peptide sequencing. Furthermore, as this strategy preserves intact protein structure, further characterization can be performed on proteins producing mass fingerprint spectra and fragmentation spectra that did not result in database protein identifications. The coupling of nonporous RP-HPLC with SDS-PAGE provides a useful alternative to two-dimensional PAGE (2-D-PAGE) for membrane protein analysis.

Breast↗

Proteomic analysis of cerebrospinal fluid from multiple sclerosis patients.

Multiple sclerosis is an autoimmune inflammatory demyelinating disease of the central nervous system. Disease mechanisms in multiple sclerosis at the molecular level remain poorly understood and no reliable proteinaceous disease markers are available yet. The goal of the present study is the construction of a protein database of two-dimensional gel electrophoresis (2-DE) separated cerebrospinal fluid (CSF) proteins from multiple sclerosis patients. By means of liquid chromatography tandem mass spectrometry 65 different proteins were identified from 300 spots. Eighteen of these proteins have not been reported previously on 2-DE gels of CSF. Here we report on the identification of these proteins and discuss their potential relation to multiple sclerosis.

Amino Acid Sequence↗

Cloning and sequence analysis of an Escherichia coli gene conferring bicyclomycin resistance.

We have cloned and sequenced DNA from Escherichia coli that, when present in a high-copy-number plasmid, confers resistance to the diketopiperazine antibiotic, bicyclomycin (Bc). The DNA includes a 378-amino-acid open reading frame (ORF), disruption of which results in the loss of Bc resistance. This ORF contains the BcR gene. Studies using the minicell expression system reveal that a polypeptide of 31 kDa is produced from this cloned region. The ORF maps at 47.1 min on the E. coli genome map. Sequence comparison between the translated ORF and a protein database reveal between 26.5 and 23.4% aa sequence homology to bacterial transmembrane (TM) proteins including those mediating chloramphenicol (Cm) and tetracycline (Tc) resistance and an arabinose-proton symport protein. Sequence analysis using the Diagon program showed the BcR gene product (BcR) had homology with the N-terminal regions of the CmR and TcR-encoded proteins and weak N-terminal homology with the arabinose-proton symport protein. Hydropathy profiles of the BcR protein and CmR products show a striking similarity, both having twelve predicted TM domains.

Amino Acid Sequence↗

Antibodies against human muscle enolase recognize a 45-kDa bacterial cell wall outer membrane enolase-like protein.

Enolase, is a glycolytic enzyme ubiquitous in higher organisms, where it forms tissue specific dimers of isoforms, also found in the cytoplasm of fermentative bacteria. The aim of this work was to identify enolase-like proteins in the cell wall of some Gram-negative bacteria using antibodies against human beta-enolase, an isoenzyme specific to skeletal and heart muscles. Cell wall outer membrane protein (OMP) preparations were obtained from 9 strains of Enterobacteriaceae and one of Pseudomonas aeruginosa. Specific enzymatic enolase activity was detected in the supernatant fractions of cytosolic and inner membrane material, but not in purified OMP preparations. Rabbit polyclonal antibodies specific against human beta-enolase were prepared and purified using immobilized human beta-enolase in affinity chromatography. In SDS-polyacrylamide gel electrophoresis and immunoblotting assay of purified OMP preparations, rabbit anti-enolase antibody interacted specifically with a few OMPs, of which a 45-kDa band also interacted with human sera of patients presenting Buerger disease and atherosclerosis. The most distinct interaction of human sera was observed with a 45-kDa OMP of Klebsiella pneumoniae. This protein was further isolated from K. pneumoniae cell mass in two ways, namely preparative SDS-polyacrylamide gel electrophoresis and specific affinity chromatography using immobilized affinity-purified rabbit antibody raised against human beta-enolase. The data obtained from tandem mass spectrometry tryptic peptide analysis and sequence comparison of human and bacterial enolases using protein databases, could reveal the similarity in the epitopes between membrane enolase-like protein from Klebsiella and human beta-enolase. The results show that the protein present in all studied strains has a common epitope on human beta-enolase. These data raise the question whether such a bacterial protein might be a marker for detecting and monitoring damage to skeletal and heart muscles.

Animals↗

Molecular characterization of a Penicillium chrysogenum exo-1,5-alpha-L-arabinanase that is structurally distinct from other arabinan-degrading enzymes.

The nucleotide sequence of the abnx cDNA gene, which encodes an exo-arabinanase (Abnx) of Penicillium chrysogenum 31B, was determined. Abnx was found to be structurally distinct from known arabinan-degrading enzymes based on its amino acid sequence and a hydrophobic cluster analysis. The protein in the protein database with the highest similarity to Abnx was the Neurospora crassa conserved hypothetical protein. The abnx cDNA gene product expressed in Escherichia coli catalyzed the release of arabinobiose from alpha-1,5-L-arabinan. The activity of the recombinant Abnx towards a series of arabino-oligosaccharides, as expressed by k(cat)/K(m) value, was greatest with arabinohexaose.

Amino Acid Sequence↗

Toward plasma proteome profiling with ion mobility-mass spectrometry.

Differential, functional, and mapping proteomic analyses of complex biological mixtures suffer from a lack of component resolution. Here we describe the application of ion mobility-mass spectrometry (IMS-MS) to this problem. With this approach, components that are separated by liquid chromatography are dispersed based on differences in their mobilities through a buffer gas prior to being analyzed by MS. The inclusion of the gas-phase dispersion provides more than an order of magnitude enhancement in component resolution at no cost to data acquisition time. Additionally, the mobility separation often removes high-abundance species from spectral regions containing low-abundance species, effectively increasing measurement sensitivity and dynamic range. Finally, collision-induced dissociation of all ions can be recorded in a single experimental sequence while conventional MS methods sequentially select precursors. The approach is demonstrated in a single, rapid (3.3 h) analysis of a plasma digest sample where abundant proteins have not been removed. Protein database searches have yielded 731 high confidence peptide assignments corresponding to 438 unique proteins. Results have been compiled into an initial analytical map to be used -after further augmentation and refinement- for comparative plasma profiling studies.

Blood Proteins↗

Isolation of a candidate gene for Norrie disease by positional cloning.

The gene for Norrie disease, an X-linked disorder characterized by progressive atrophy of the eyes, mental disturbances and deafness, has been mapped to chromosome Xp11.4 close to DXS7 and the monoamine oxidase (MAO) genes. By subcloning a YAC with a 640 kilobases (kb) insert which spans the DXS7-MAOB interval we have generated a cosmid contig which extends 250 kb beyond the MAOB gene. With one of these cosmids, microdeletions were detected in several patients with Norrie disease. Screening of cDNA libraries has enabled us to isolate and sequence a likely candidate gene for Norrie disease which is expressed in retina, choroid and fetal brain. No homologous sequences were found in DNA and protein databases indicating that this cDNA is part of a gene encoding a 'pioneer' protein.

Adult↗

A two-dimensional electrophoresis proteomic reference map and systematic identification of 1367 proteins from a cell suspension culture of the model legume Medicago truncatula.

The proteome of a Medicago truncatula cell suspension culture was analyzed using two-dimensional electrophoresis and nanoscale HPLC coupled to a tandem Q-TOF mass spectrometer (QSTAR Pulsar i) to yield an extensive protein reference map. Coomassie Brilliant Blue R-250 was used to visualize more than 1661 proteins, which were excised, subjected to in-gel trypsin digestion, and analyzed using nanoscale HPLC/MS/MS. The resulting spectral data were queried against a custom legume protein database using the MASCOT search engine. A total of 1367 of the 1661 proteins were identified with high rigor, yielding an identification success rate of 83% and 907 unique protein accession numbers. Functional annotation of the M. truncatula suspension cell proteins revealed a complete tricarboxylic acid cycle, a nearly complete glycolytic pathway, a significant portion of the ubiquitin pathway with the associated proteolytic and regulatory complexes, and many enzymes involved in secondary metabolism such as flavonoid/isoflavonoid, chalcone, and lignin biosynthesis. Proteins were also identified from most other functional classes including primary metabolism, energy production, disease/defense, protein destination/storage, protein synthesis, transcription, cell growth/division, and signal transduction. This work represents the most extensive proteomic description of M. truncatula suspension cells to date and provides a reference map for future comparative proteomic and functional genomic studies of the response of these cells to biotic and abiotic stress.

Amino Acid Sequence↗

The complete nucleotide sequence of the rice grassy stunt virus genome and genomic comparisons with viruses of the genus Tenuivirus.

Rice grassy stunt virus (RGSV, IRRI isolate) has six genomic RNA segments. The nucleotide (nt) sequences of RNAs 1-4 were determined. The cumulative length of the RGSV genome, including RNAs 5 and 6, was 25142 nt. All six RNA segments had an ambisense coding strategy and almost identical terminal sequences over 17 nt. The virus complementary (vc) sequence of the largest segment, RNA1, had an open reading frame encoding a protein of Mr 339133 (the 339.1K protein), while the virus sense (v) sequence encoded a protein of Mr 18910 in the 5'-proximal region. The predicted 339.1K protein contained the highly conserved motifs of the RNA-dependent RNA polymerase and a short but distinct Arg/Gly-rich stretch at the C terminus. The putative RNA polymerase showed strong similarity with that of rice stripe tenuivirus (RSV); they shared 37.9% amino acid identity over 2140 residues. The predicted proteins of Mr 23280 on vRNA2 and 93 879 on vcRNA2 were only slightly similar in sequence to the proteins encoded by vRNA2 and vcRNA2 of other tenuiviruses. The predicted proteins encoded by RNA3 and RNA4 did not show significant similarity to any database proteins. Only the putative RNA polymerase encoded on RNA1 was well-conserved between RGSV and RSV. The low sequence similarities in proteins encoded by RNAs 2, 5 and 6, together with the unique RNA segments 3 and 4, indicate that RGSV may be distinct from other tenuiviruses.

Base Sequence↗

Proteome analysis of the effect of mucoid conversion on global protein expression in Pseudomonas aeruginosa strain PAO1 shows induction of the disulfide bond isomerase, dsbA.

Pseudomonas aeruginosa strains that cause chronic pulmonary infections in cystic fibrosis patients typically undergo mucoid conversion. The mucoid phenotype indicates alginate overproduction and is often due to defects in MucA, an antisigma factor that controls the activity of sigma-22 (AlgT [also called AlgU]), which is required for the activation of genes for alginate biosynthesis. In this study we hypothesized that mucoid conversion may be part of a larger response that activates genes other than those for alginate synthesis. To address this, a two-dimensional (2-D) gel analysis was employed to compare total proteins in strain PAO1 to those of its mucA22 derivative, PDO300, in order to identify protein levels enhanced by mucoid conversion. Six proteins that were clearly more abundant in the mucoid strain were observed. The amino termini of such proteins were determined and used to identify the gene products in the genomic database. Proteins involved in alginate biosynthesis were expected among these, and two (AlgA and AlgD) were identified. This result verified that the 2-D gel approach could identify gene products under sigma-22 control and upregulated by mucA mutation. Two other protein spots were also clearly upregulated in the mucA22 background, and these were identified as porin F (an outer membrane protein) and a homologue of DsbA (a disulfide bond isomerase). Single-copy gene fusions were constructed to test whether these proteins were enhanced in the mucoid strain due to increased transcription. The oprF-lacZ fusion showed little difference in levels of expression in the two strains. However, the dsbA-lacZ fusion showed two- to threefold higher expression in PDO300 than in PAO1, suggesting that its promoter was upregulated by the deregulation of sigma-22 activity. A dsbA-null mutant was constructed in PAO1 and shown to have defects predicted for a cell with reduced disulfide bond isomerase activity, namely, reduction in periplasmic alkaline phosphatase activity, increased sensitivity to dithiothreitol, reduced type IV pilin-mediated twitching motility, and reduced accumulation of extracellular proteases, including elastase. Although efficient secretion of elastase in the dsbA mutant was still demonstrable, the elastase produced appeared to be unstable, possibly as a result of mispaired disulfide bonds. Disruption of dsbA in the mucoid PDO300 background did not affect alginate production. Thus, even though dsbA is coregulated with mucoid conversion, it was not required for alginate production. This suggests that mucA mutation, which deregulates sigma-22, results in a global response that includes other factors in addition to increasing the production of alginate.

Adenosine Triphosphatases↗

Proteomic identification of 14-3-3zeta as a mitogen-activated protein kinase-activated protein kinase 2 substrate: role in dimer formation and ligand binding.

Mitogen-activated protein kinase (MAPK)-activated protein kinase 2 (MAPKAPK2) mediates multiple p38 MAPK-dependent inflammatory responses. To define the signal transduction pathways activated by MAPKAPK2, we identified potential MAPKAPK2 substrates by using a functional proteomic approach consisting of in vitro phosphorylation of neutrophil lysate by active recombinant MAPKAPK2, protein separation by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE), and phosphoprotein identification by peptide mass fingerprinting with matrix-assisted laser desorption ionization mass spectrometry (MALDI-MS) and protein database analysis. One of the eight candidate MAPKAPK2 substrates identified was the adaptor protein, 14-3-3zeta. We confirmed that MAPKAPK2 interacted with and phosphorylated 14-3-3zeta in vitro and in HEK293 cells. The chemoattractant formyl-methionyl-leucyl-phenylalanine (fMLP) stimulated p38-MAPK-dependent phosphorylation of 14-3-3 proteins in human neutrophils. Mutation analysis showed that MAPKAPK2 phosphorylated 14-3-3zeta at Ser-58. Computational modeling and calculation of theoretical binding energies predicted that both phosphorylation at Ser-58 and mutation of Ser-58 to Asp (S58D) compromised the ability of 14-3-3zeta to dimerize. Experimentally, S58D mutation significantly impaired both 14-3-3zeta dimerization and binding to Raf-1. These data suggest that MAPKAPK2-mediated phosphorylation regulates 14-3-3zeta functions, and this MAPKAPK2 activity may represent a novel pathway mediating p38 MAPK-dependent inflammation.

14-3-3 Proteins↗

Computational identification of strain-, species- and genus-specific proteins.

BACKGROUND: The identification of unique proteins at different taxonomic levels has both scientific and practical value. Strain-, species- and genus-specific proteins can provide insight into the criteria that define an organism and its relationship with close relatives. Such proteins can also serve as taxon-specific diagnostic targets. DESCRIPTION: A pipeline using a combination of computational and manual analyses of BLAST results was developed to identify strain-, species-, and genus-specific proteins and to catalog the closest sequenced relative for each protein in a proteome. Proteins encoded by a given strain are preliminarily considered to be unique if BLAST, using a comprehensive protein database, fails to retrieve (with an e-value better than 0.001) any protein not encoded by the query strain, species or genus (for strain-, species- and genus-specific proteins respectively), or if BLAST, using the best hit as the query (reverse BLAST), does not retrieve the initial query protein. Results are manually inspected for homology if the initial query is retrieved in the reverse BLAST but is not the best hit. Sequences unlikely to retrieve homologs using the default BLOSUM62 matrix (usually short sequences) are re-tested using the PAM30 matrix, thereby increasing the number of retrieved homologs and increasing the stringency of the search for unique proteins. The above protocol was used to examine several food- and water-borne pathogens. We find that the reverse BLAST step filters out about 22% of proteins with homologs that would otherwise be considered unique at the genus and species levels. Analysis of the annotations of unique proteins reveals that many are remnants of prophage proteins, or may be involved in virulence. The data generated from this study can be accessed and further evaluated from the CUPID (Core and Unique Protein Identification) system web site (updated semi-annually) at http://pir.georgetown.edu/cupid. CONCLUSION: CUPID provides a set of proteins specific to a genus, species or a strain, and identifies the most closely related organism.

Algorithms↗

Application of string kernels in protein sequence classification.

INTRODUCTION: The production of biological information has become much greater than its consumption. The key issue now is how to organise and manage the huge amount of novel information to facilitate access to this useful and important biological information. One core problem in classifying biological information is the annotation of new protein sequences with structural and functional features. METHOD: This article introduces the application of string kernels in classifying protein sequences into homogeneous families. A string kernel approach used in conjunction with support vector machines has been shown to achieve good performance in text categorisation tasks. We evaluated and analysed the performance of this approach, and we present experimental results on three selected families from the SCOP (Structural Classification of Proteins) database. We then compared the overall performance of this method with the existing protein classification methods on benchmark SCOP datasets. RESULTS: According to the F1 performance measure and the rate of false positive (RFP) measure, the string kernel method performs well in classifying protein sequences. The method outperformed all the generative-based methods and is comparable with the SVM-Fisher method. DISCUSSION: Although the string kernel approach makes no use of prior biological knowledge, it still captures sufficient biological information to enable it to outperform some of the state-of-the-art methods.

Algorithms↗

A group of expressed cDNA sequences from the wheat fungal leaf blotch pathogen, Mycosphaerella graminicola (Septoria tritici).

A group of expressed sequence tags (ESTs) from the wheat fungal pathogen Mycosphaerella graminicola utilizing ammonium as a nitrogen source has been analyzed. Single pass sequences of complementary DNAs from 986 clones were determined. Contig analysis and sequence comparisons allowed 704 unique ESTs (unigenes) to be identified, of which 148 appeared as multiple copies. Searches of the nrdb95 protein database at EMBL using the BLAST2x algorithm revealed 407 (57.8%) sequences that generated high to moderate high scoring pairs with proteins of known and unknown function. The rest of the sequences (297) showed either weak or no similarities to database entries. Among the unigenes with assigned function, 26.7% were involved in primary metabolism and 17.9% were associated with protein and RNA metabolism. Fewer clones were ascribed roles in signal transduction (4.9%), transport and secretion (6.1%), cell structure (3.1%), and cell division (3.6%). Approximately 18.1% of the identities found were to hypothetical or unknown proteins mainly from the yeasts Saccharomyces cerevisiae and Schizosaccaromyces pombe. Comparison of the 297 sequences with no clear function to other fungal ESTs in the public domain revealed 12 sequences that had high to moderate similarity to Neurospora crassa, Emericella (Aspergillus) nidulans, or Magnaporthe grisea sequences.

Ascomycota↗