PubMed HealthSearch

Biomedical subjects

D Frishman

Publications and source records attributed to D Frishman.

17 recordsLinked to original sources

Combining diverse evidence for gene recognition in completely sequenced bacterial genomes.

Analysis of a newly sequenced bacterial genome starts with identification of protein-coding genes. Functional assignment of proteins requires the exact knowledge of protein N-termini. We present a new program ORPHEUS that identifies candidate genes and accurately predicts gene starts. The analysis starts with a database similarity search and identification of reliable gene fragments. The latter are used to derive statistical characteristics of protein-coding regions and ribosome-binding sites and to predict the complete set of genes in the analyzed genome. In a test on Bacillus subtilis and Escherichia coli genomes, the program correctly identified 93.3% (resp. 96.3%) of experimentally annotated genes longer than 100 codons described in the PIR-International database, and for these genes 96.3% (83.9%) of starts were predicted exactly. Furthermore, 98.9% (99.1%) of genes longer than 100 codons annotated in GenBank were found, and 92.9% (75.7%) of predicted starts coincided with the feature table description. Finally, for the complete gene complements of B.subtilis and E.coli , including genes shorter than 100 codons, gene prediction accuracy was 88.9 and 87.1%, respectively, with 94.2 and 76.7% starts coinciding with the existing annotation.

Algorithms

Iron-regulatory protein-1 (IRP-1) is highly conserved in two invertebrate species--characterization of IRP-1 homologues in Drosophila melanogaster and Caenorhabditis elegans.

Iron-regulatory protein-1 (IRP-1) plays a dual role as a regulatory RNA-binding protein and as a cytoplasmic aconitase. When bound to iron-responsive elements (IRE), IRP-1 post-transcriptionally regulates the expression of mRNAs involved in iron metabolism. IRP have been cloned from several vertebrate species. Using a degenerate-primer PCR strategy and the screening of data bases, we now identify the homologues of IRP-1 in two invertebrate species, Drosophila melanogaster and Caenorhabditis elegans. Comparative sequence analysis shows that these invertebrate IRP are closely related to vertebrate IRP, and that the amino acid residues that have been implicated in aconitase function are particularly highly conserved, suggesting that invertebrate IRP may function as cytoplasmic aconitases. Antibodies raised against recombinant human IRP-1 immunoprecipitate the Drosophila homologue expressed from the cloned cDNA. In contrast to vertebrates, two IRP-1 homologues (Drosophila IRP-1A and Drosophila IRP-1B), displaying 86% identity to each other, are expressed in D. melanogaster. Both of these homologues are distinct from vertebrate IRP-2. In contrast to the mammalian system where the two IRP (IRP-1 and IRP-2) are differentially expressed, Drosophila IRP-1A and Drosophila IRP-1B are not preferentially expressed in specific organs. The localization of Drosophila IRP-1A to position 94C1-8 and of Drosophila IRP-1B to position 86B3-6 on the right arm of chromosome 3 and the availability of an IRP-1 cDNA from C. elegans will facilitate a genetic analysis of the IRE/IRP system, thus opening a new avenue to explore this regulatory network.

Amino Acid Sequence

MIPS: a database for protein sequences and complete genomes.

The MIPS group [Munich Information Center for Protein Sequences of the German National Center for Environment and Health (GSF)] at the Max-Planck-Institute for Biochemistry, Martinsried near Munich, Germany, is involved in a number of data collection activities, including a comprehensive database of the yeast genome, a database reflecting the progress in sequencing the Arabidopsis thaliana genome, the systematic analysis of other small genomes and the collection of protein sequence data within the framework of the PIR-International Protein Sequence Database (described elsewhere in this volume). Through its WWW server (http://www.mips.biochem.mpg.de ) MIPS provides access to a variety of generic databases, including a database of protein families as well as automatically generated data by the systematic application of sequence analysis algorithms. The yeast genome sequence and its related information was also compiled on CD-ROM to provide dynamic interactive access to the 16 chromosomes of the first eukaryotic genome unraveled.

Amino Acid Sequence

Single-read sequence tags of a limited number of genomic DNA fragments provide an inexpensive tool for comparative genome analysis.

Single-read sequences from both ends of 415 3-kb average size genomic DNA fragments of Candida albicans were compared with the complete sequence data of Saccharomyces cerevisiae. Comparison at the protein level, translated DNA against protein sequences, revealed 138 sequence tags with clear similarity to S. cerevisiae proteins or open reading frames. One case of synteny was found for the open reading frames of RAD16 and LYS2, which are adjacent to each other in S. cerevisiae and C. albicans.

Adenosine Triphosphatases

Comprehensive, comprehensible, distributed and intelligent databases: current status.

MOTIVATION: It is only a matter of time until a user will see not many but one integrated database of information for molecular biology. Is this true? Is it a good thing? Why will it happen? Where are we now? What developments are fostering and what developments are impeding progress towards this end? SUPPLEMENTARY INFORMATION: A list of WWW resources devoted to database issues in molecular biology is available at http://www.mips.biochem.mpg.de CONTACT: frishman@mips.biochem.mpg.de

Computational Biology

Variations of the C2H2 zinc finger motif in the yeast genome and classification of yeast zinc finger proteins.

The PROSITE pattern Zinc_Finger_C2H2 was extended to permit the detection of all C2H2 zinc fingers and their parent proteins in the recently completed sequence of the yeast genome. Additionally, a new computer program was written that extracts other zinc binding motifs (non C2H2 'fingers'), overlapping with the classical zinc finger pattern, from the found set of yeast C2H2 fingers. The complete and correct detection of all fingers is a prerequisite for the classification of the yeast zinc finger proteins in functional terms. The detected 53 yeast C2H2 zinc finger proteins do not contain finger clusters with 10 or more repeats, as is frequently found in higher eukaryotes. Only three proteins contain four or more fingers in a cluster. Moreover, nearly all 27 yeast proteins with tandem arrays of two or three finger domains can be classified into nine subgroups with high sequence conservation in their finger clusters, in particular of their DNA recognition helices. These results and application of the recently elaborated finger/DNA recognition rules suggest that the yeast proteins belonging to the same subgroup may recognize identical or very similar DNA sites.

Amino Acid Sequence

Overview of the yeast genome.

The collaboration of more than 600 scientists from over 100 laboratories to sequence the Saccharomyces cerevisiae genome was the largest decentralised experiment in modern molecular biology and resulted in a unique data resource representing the first complete set of genes from a eukaryotic organism. 12 million bases were sequenced in a truly international effort involving European, US, Canadian and Japanese laboratories. While the yeast genome represents only a small fraction of the information in today's public sequence databases, the complete, ordered and non-redundant sequence provides an invaluable resource for the detailed analysis of cellular gene function and genome architecture. In terms of throughput, completeness and information content, yeast has always been the lead eukaryotic organism in genomics; it is still the largest genome to be completely sequenced.

Chromosome Mapping

Seventy-five percent accuracy in protein secondary structure prediction.

In this study we present an accurate secondary structure prediction procedure by using an query and related sequences. The most novel aspect of our approach is its reliance on local pairwise alignment of the sequence to be predicted with each related sequence rather than utilization of a multiple alignment. The residue-by-residue accuracy of the method is 75% in three structural states after jack-knife tests. The gain in prediction accuracy compared with the existing techniques, which are at best 72%, is achieved by secondary structure propensities based on both local and long-range effects, utilization of similar sequence information in the form of carefully selected pairwise alignment fragments, and reliance on a large collection of known protein primary structures. The method is especially appropriate for large-scale sequence analysis of efforts such as genome characterization, where precise and significant multiple sequence alignments are not available or achievable.

Algorithms

The future of protein secondary structure prediction accuracy.

BACKGROUND: The accuracy of secondary structure prediction for a protein from knowledge of its sequence has been significantly improved by about 7% to the 70-75% range by inclusion of information residing in sequences similar to the query sequence. The scientific literature has been inconsistent, if not negative, regarding chances for further improvement from the vast knowledge to be provided by genome sequencing efforts. RESULTS: By applying a prediction technique that is particularly sensitive to added sequence information to a standard set of query sequences with related primary structures taken from chronologically successive releases of the SWISS-PROT database, it is shown that prediction accuracy can be expected to reach 80-85% with a large 10-fold increase in present sequence knowledge. CONCLUSIONS: Even with present prediction approaches, improvement in prediction accuracy can still be expected, albeit limited to no more than 10%.

Algorithms

Protein structural classes in five complete genomes.

The predicted distribution of globular proteins over folding types in five complete genomes differs from the tendencies observed in known protein structures. The ratio between the number of predicted membrane and globular proteins is conserved.

Bacterial Proteins

Conservation of aconitase residues revealed by multiple sequence analysis. Implications for structure/function relationships.

Aconitases have recently regained much attention, because one member of this family, iron regulatory protein-1 (IRP-1), has been found to play a dual role as a cytoplasmic aconitase and a regulatory RNA-binding protein. This finding has highlighted a novel role for Fe-S clusters as post-translational regulatory switches. We have aligned 28 members of the Fe-S isomerase family, identified highly conserved amino acid residues, and integrated this information with data on the crystallographic structure of mammalian mitochondrial aconitase. We propose structural and/or functional roles for the previously unrecognized conserved residues. Our findings illustrate the value of detailed protein sequence analysis when high-resolution crystallographic data are already available.

Aconitate Hydratase

DSBC protein: a new member of the thioredoxin fold-containing family.

Prediction of the DsbC protein secondary structure has been performed using a novel prediction technique which is based on consideration of both local and long-range interactions between amino acid residues. The C-terminal portion of the protein is shown to contain the thioredoxin folding motif. The N-terminal part represents a yet unknown structural domain.

Amino Acid Sequence

Incorporation of non-local interactions in protein secondary structure prediction from the amino acid sequence.

Existing approaches to protein secondary structure prediction from the amino acid sequence usually rely on the statistics of local residue interactions within a sliding window and the secondary structural state of the central residue. The practically achieved accuracy limit of such single residue and single sequence prediction methods is 65% in three structural stages (alpha-helix, beta-strand and coil). Further improvement in the prediction quality is likely to require exploitation of various aspects of three-dimensional protein architecture. Here we make such an attempt and present an accurate algorithm for secondary structure prediction based on recognition of potentially hydrogen-bonded residues in a single amino acid sequence. The unique feature of our approach involves database-derived statistics on residue type occurrences in different classes of beta-bridges to delineate interacting beta-strands. The alpha-helical structures are also recognized on the basis of amino acid occurrences in hydrogen-bonded pairs (i,i + 4). The algorithm has a prediction accuracy of 68% in three structural stages, relies only on a single protein sequence as input and has the potential to be improved by 5-7% if homologous aligned sequences are also considered.

Algorithms

Knowledge-based protein secondary structure assignment.

We have developed an automatic algorithm STRIDE for protein secondary structure assignment from atomic coordinates based on the combined use of hydrogen bond energy and statistically derived backbone torsional angle information. Parameters of the pattern recognition procedure were optimized using designations provided by the crystallographers as a standard-of-truth. Comparison to the currently most widely used technique DSSP by Kabsch and Sander (Biopolymers 22:2577-2637, 1983) shows that STRIDE and DSSP assign secondary structural states in 58 and 31% of 226 protein chains in our data sample, respectively, in greater agreement with the specific residue-by-residue definitions provided by the discoverers of the structures while in 11% of the chains, the assignments are the same. STRIDE delineates every 11th helix and every 32nd strand more in accord with published assignments.

Algorithms

Recognition of distantly related proteins through energy calculations.

A new method to detect remote relationships between protein sequences and known three-dimensional structures based on direct energy calculations and without reliance on statistics has been developed. The likelihood of a residue to occupy a given position on the structural template was represented by an estimate of the stabilization free energy made after explicit prediction of the substituted side chain conformation. The profile matrix derived from these energy values and modified by increasing the residue self-exchange values successfully predicted compatibility of heat-shock protein and globin sequences with the three-dimensional structures of actin and phycocyanin, respectively, from a full protein sequence databank search. The high sensitivity of the method makes it a unique tool for predicting the three-dimensional fold for the rapidly growing number of protein sequences.

Amino Acid Sequence

Recognition of distantly related protein sequences using conserved motifs and neural networks.

A sensitive technique for protein sequence motif recognition based on neural networks has been developed. It involves three major steps. (1) At each appropriate alignment position of a set of N matched sequences, a set of N aligned oligopeptides is specified with preselected window length. N neural nets are subsequently and successively trained on N-1 amino acid spans after eliminating each ith oligopeptide. A test for recognition of each of the ith spans is performed. The average neural net recognition over N such trials is used as a measure of conservation for the particular windowed region of the multiple alignment. This process is repeated for all possible spans of given length in the multiple alignment. (2) The M most conserved regions are regarded as motifs and the oligopeptides within each are used to train intensively M individual neural networks. (3) The M networks are then applied in a search for related primary structures in a databank of known protein sequences. The oligopeptide spans in the database sequence with strongest neural net output for each of the M networks are saved and then scored according to the output signals and the proper combination that follows the expected N- to C-terminal sequence order. The motifs from the database with highest similarity scores can then be used to retrain the M neural nets, which can be subsequently utilized for further searches in the databank, thus providing even greater sensitivity to recognize distant familial proteins. This technique was successfully applied to the integrase, DNA-polymerase and immunoglobulin families.

Aldehyde Dehydrogenase

Hairpieces.

Explore the source record for details and available documents.

Female