PubMed Health⌕ Search

Biomedical subjects

H Hegyi

Publications and source records attributed to H Hegyi.

17 recordsLinked to original sources

Annotation transfer for genomics: measuring functional divergence in multi-domain proteins.

Annotation transfer is a principal process in genome annotation. It involves "transferring" structural and functional annotation to uncharacterized open reading frames (ORFs) in a newly completed genome from experimentally characterized proteins similar in sequence. To prevent errors in genome annotation, it is important that this process be robust and statistically well-characterized, especially with regard to how it depends on the degree of sequence similarity. Previously, we and others have analyzed annotation transfer in single-domain proteins. Multi-domain proteins, which make up the bulk of the ORFs in eukaryotic genomes, present more complex issues in functional conservation. Here we present a large-scale survey of annotation transfer in these proteins, using scop superfamilies to define domain folds and a thesaurus based on SWISS-PROT keywords to define functional categories. Our survey reveals that multi-domain proteins have significantly less functional conservation than single-domain ones, except when they share the exact same combination of domain folds. In particular, we find that for multi-domain proteins, approximate function can be accurately transferred with only 35% certainty for pairs of proteins sharing one structural superfamily. In contrast, this value is 67% for pairs of single-domain proteins sharing the same structural superfamily. On the other hand, if two multi-domain proteins contain the same combination of two structural superfamilies the probability of their sharing the same function increases to 80% in the case of complete coverage along the full length of both proteins, this value increases further to > 90%. Moreover, we found that only 70 of the current total of 455 structural superfamilies are found in both single and multi-domain proteins and only 14 of these were associated with the same function in both categories of proteins. We also investigated the degree to which function could be transferred between pairs of multi-domain proteins with respect to the degree of sequence similarity between them, finding that functional divergence at a given amount of sequence similarity is always about two-fold greater for pairs of multi-domain proteins (sharing similarity over a single domain) in comparison to pairs of single-domain ones, though the overall shape of the relationship is quite similar. Further information is available at http://partslist.org/func or http://bioinfo.mbb.yale.edu/partslist/func.

Computational Biology↗

Protein folds in the worm genome.

We survey the protein folds in the worm genome, using pairwise and multiple-sequence comparison methods (i.e. FASTA and PSI-blast). Overall, we find that approximately 250 folds match approximately 8000 domains in approximately 4500 ORFs, about 32 matches per fold involving a quarter of the total worm ORFs. We compare the folds in the worm genome to those in other model organisms, in particular yeast and E. coli, and find that the worm shares more folds with the phylogenetically closer yeast than with E. coli. There appear to be 36 folds unique to the worm compared to these two model organisms, and many of these are obviously implicated in aspects of multicellularity. The most common fold in the worm genome is the immunoglobulin fold, and many of the common folds are repeated in various combinations and permutations in multidomain proteins. In addition, an approach is presented for the identification of "sure" and "marginal" membrane proteins. When applied to the worm genome, this reveals a much greater relative prevalence of proteins with seven transmembrane helices in comparison to the other completely sequenced genomes, which are not of metazoans. Combining these analyses with some other simple filters allows one to identify ORFs that potentially code for soluble proteins of unknown fold, which may be promising targets for experimental investigation in structural genomics. A regularly updated worm fold analysis will be available from bioinfo.mbb.yale.edu/genome/worm.

Animals↗

Genome analyses of spirochetes: a study of the protein structures, functions and metabolic pathways in Treponema pallidum and Borrelia burgdorferi.

We perform a comprehensive genome analysis on two spirochetes, T. pallidum and B. burgdorferi. First, we focus on the occurrence of protein structures in these organisms. We find that there are only a few spirochete-specific folds, relative to those in other types of bacteria. The most common fold, by far, in the spirochetes is the P-loop NTP hydrolase, followed by the TIM barrel. These folds also happen to be amongst the most multifunctional of the known folds. We also survey the membrane-protein structures in T. pallidum and find a notable large family with twelve transmembrane (TM) helices, reflecting the prevalence of 12-TM transporters in bacteria. Then we move to analysis of the metabolic pathways and overall metabolism in the spirochetes, using the metabolic-flux-balancing method. We find that the lipid biosynthesis pathway is absent from the spirochetes. This strongly limits the degree to which these organisms can metabolize NADPH. In turn, we find that the spirochetes distribute flux disproportionately through the glycolytic pathway instead of the NADPH-providing pentose phosphate pathway. Further information is available at http://bioinfo.mbb.yale.edu

Bacterial Proteins↗

The relationship between protein structure and function: a comprehensive survey with application to the yeast genome.

For most proteins in the genome databases, function is predicted via sequence comparison. In spite of the popularity of this approach, the extent to which it can be reliably applied is unknown. We address this issue by systematically investigating the relationship between protein function and structure. We focus initially on enzymes functionally classified by the Enzyme Commission (EC) and relate these to by structurally classified domains the SCOP database. We find that the major SCOP fold classes have different propensities to carry out certain broad categories of functions. For instance, alpha/beta folds are disproportionately associated with enzymes, especially transferases and hydrolases, and all-alpha and small folds with non-enzymes, while alpha+beta folds have an equal tendency either way. These observations for the database overall are largely true for specific genomes. We focus, in particular, on yeast, analyzing it with many classifications in addition to SCOP and EC (i.e. COGs, CATH, MIPS), and find clear tendencies for fold-function association, across a broad spectrum of functions. Analysis with the COGs scheme also suggests that the functions of the most ancient proteins are more evenly distributed among different structural classes than those of more modern ones. For the database overall, we identify the most versatile functions, i.e. those that are associated with the most folds, and the most versatile folds, associated with the most functions. The two most versatile enzymatic functions (hydro-lyases and O-glycosyl glucosidases) are associated with seven folds each. The five most versatile folds (TIM-barrel, Rossmann, ferredoxin, alpha-beta hydrolase, and P-loop NTP hydrolase) are all mixed alpha-beta structures. They stand out as generic scaffolds, accommodating from six to as many as 16 functions (for the exceptional TIM-barrel). At the conclusion of our analysis we are able to construct a graph giving the chance that a functional annotation can be reliably transferred at different degrees of sequence and structural similarity. Supplemental information is available from http://bioinfo.mbb.yale.edu/genome/foldfunc++ +.

Enzymes↗

The domain-server: direct prediction of protein domain-homologies from BLAST search.

RESULTS: A WWW server for protein domain homology prediction, based on BLAST search and a simple data-mining algorithm (Hegyi,H. and Pongor,S. (1993) Comput. Appl. Biosci., 9, 371-372), was constructed providing a tabulated list and a graphic plot of similarities. AVAILABILITY: http://www.icgeb.trieste.it/domain. Mirror site is available at http://sbase.abc.hu/domain. A standalone programme will be available on request. SUPPLEMENTARY INFORMATION: A series of help files is available at the above addresses.

Algorithms↗

Comparing genomes in terms of protein structure: surveys of a finite parts list.

We give an overview of the emerging field of structural genomics, describing how genomes can be compared in terms of protein structure. As the number of genes in a genome and the total number of protein folds are both quite limited, these comparisons take the form of surveys of a finite parts list, similar in respects to demographic censuses. Fold surveys have many similarities with other whole-genome characterizations, e.g., analyses of motifs or pathways. However, structure has a number of aspects that make it particularly suitable for comparing genomes, namely the way it allows for the precise definition of a basic protein module and the fact that it has a better defined relationship to sequence similarity than does protein function. An essential requirement for a structure survey is a library of folds, which groups the known structures into 'fold families.' This library can be built up automatically using a structure comparison program, and we described how important objective statistical measures are for assessing similarities within the library and between the library and genome sequences. After building the library, one can use it to count the number of folds in genomes, expressing the results in the form of Venn diagrams and 'top-10' statistics for shared and common folds. Depending on the counting methodology employed, these statistics can reflect different aspects of the genome, such as the amount of internal duplication or gene expression. Previous analyses have shown that the common folds shared between very different microorganisms, i.e., in different kingdoms, have a remarkably similar structure, being comprised of repeated strand-helix-strand super-secondary structure units. A major difficulty with this sort of 'fold-counting' is that only a small subset of the structures in a complete genome are currently known and this subset is prone to sampling bias. One way of overcoming biases is through structure prediction, which can be applied uniformly and comprehensively to a whole genome. Various investigators have, in fact, already applied many of the existing techniques for predicting secondary structure and transmembrane (TM) helices to the recently sequenced genomes. The results have been consistent: microbial genomes have similar fractions of strands and helices even though they have significantly different amino acid composition. The fraction of membrane proteins with a given number of TM helices falls off rapidly with more TM elements, approximately according to a Zipf law. This latter finding indicates that there is no preference for the highly studied 7-TM proteins in microbial genomes. Continuously updated tables and further information pertinent to this review are available over the web at http://bioinfo.mbb.yale.edu/genome.

Amino Acid Sequence↗

The SBASE protein domain library, release 5.0: a collection of annotated protein sequence segments.

SBASE 5.0 is the fifth release of SBASE, a collection of annotated protein domain sequences that represent various structural, functional, ligand-binding and topogenic segments of proteins. SBASE was designed to facilitate the detection of functional homologies and can be searched with standard database-search programs. The present release contains over 79863 entries provided with standardized names and is cross-referenced to all major sequence databases and sequence pattern collections. The information is assigned to individual domains rather than to entire protein sequences, thus SBASE contains substantially more cross-references and links than do the protein sequence databases. The entries are clustered into >16 000 groups in order to facilitate the detection of distant similarities. SBASE 5.0 is freely available by anonymous 'ftp' file transfer from . Automated searching of SBASE with BLAST can be carried out with the WWW-server . and with the electronic mail server which now also provides a graphic representation of the homologies. A related WWW-server and e-mail server predicts SBASE domain homologies on the basis of SWISS-PROT searches.

Amino Acid Sequence↗

On the classification and evolution of protein modules.

Our efforts to classify the functional units of many proteins, the modules, are reviewed. The data from the sequencing projects for various model organisms are extremely helpful in deducing the evolution of proteins and modules. For example, a dramatic increase of modular proteins can be observed from yeast to C. elegans in accordance with new protein functions that had to be introduced in multicellular organisms. Our sequence characterization of modules relies on sensitive similarity search algorithms and the collection of multiple sequence alignments for each module. To trace the evolution of modules and to further automate the classification, we have developed a sequence and a module alerting system that checks newly arriving sequence data for the presence of already classified modules. Using these systems, we were able to identify an unexpected similarity between extracellular C1Q modules with bacterial proteins.

Amino Acid Sequence↗

The Sequence Alerting Server--a new WEB server.

UNLABELLED: A Sequence Alerting Server with a WWW interface is described which informs users with query sequences in database searches about new entries in protein databases related to their query. AVAILABILITY: The server address is http://www.bork.embl-heidelberg.de/alerting/.

Amino Acid Sequence↗

The SBASE protein domain library, Release 4.0: a collection of annotated protein sequence segments.

SBASE 4.0 is the fourth release of SBASE, a collection of annotated protein domain sequences that represent various structural, functional, ligand binding and topogenic segments of proteins. SBASE was designed to facilitate the detection of functional homologies and can be searched with standard database search tools, such as FASTA and BLAST3. The present release contains 61 137 entries provided with standardized names and cross-referenced to all major protein, nucleic acid and sequence pattern collections. The entries are clustered into 13 155 groups in order to facilitate detection of distant similarities. SBASE 4.0 is freely available by anonymous ftp file transfer from ftp.icgeb.trieste.it. Individual records can be retrieved with the gopher server at icgeb.trieste.it and with a World Wide Web server at http://www.icgeb.trieste.it. Automated searching of SBASE with BLAST can be carried out with the electronic mail server sbase@icgeb.trieste.it, which now also provides a graphic representation of the homologies. A related mail server, domain@hubi.abc.hu, assigns SBASE domain homologies on the basis of SWISS-PROT searches.

Amino Acid Sequence↗

The protein phosphatase 2C (PP2C) superfamily: detection of bacterial homologues.

A thorough sequence analysis of the various members of the eukaryotic protein serine/threonine phosphatase 2C (PP2C) family revealed the conservation of 11 motifs. These motifs could be identified in numerous other sequences, including fungal adenylate cyclases that are predicted to contain a functionally active PP2C domain, and a family of prokaryotic serine/threonine phosphatases including SpoIIE. Phylogenetic analysis of all the proteins indicates a widespread sequence family for which a considerable number of isoenzymes can be inferred.

Amino Acid Sequence↗

Towards an intelligent system for the automatic assignment of domains in globular proteins.

The automatic identification of protein domains from coordinates is the first step in the classification of protein folds and hence is required for databases to guide structure prediction. Most algorithms encode a single concept based and sometimes do not yield assignments that are consistent with the generally accepted perception. Our development of an automatic approach to identify reliably domains from protein coordinates is described. The algorithm is benchmarked against a manual identification of the domains in 284 representative protein chains. The first step is the domain assignment by distance (DAD) algorithm that considers the density of inter-residue contacts represented in a contact matrix. The algorithm yields 85% agreement with the manual assignment. The paper then considers how the reliability of these assignments could be evaluated. Finally the use of structural comparisons using the STAMP algorithm to validate domain assignment is reported on a test case.

Algorithms↗

The SBASE protein domain library, release 3.0: a collection of annotated protein sequence segments.

SBASE 3.0 is the third release of SBASE, a collection of annotated protein domain sequences. SBASE entries represent various structural, functional, ligand-binding and topogenic segments of proteins as defined by their publishing authors. SBASE can be used for establishing domain homologies using different database-search tools such as FASTA [Lipman and Pearson (1985) Science, 227, 1436-1441], and BLAST3 [Altschul and Lipman (1990) Proc. Natl. Acad. Sci. USA, 87, 5509-5513] which is especially useful in the case of loosely defined domain types for which efficient consensus patterns can not be established. The present release contains 41,749 entries provided with standardized names and cross-referenced to the major protein and nucleic acid databanks as well as to the PROSITE catalogue of protein sequence patterns. The entries are clustered into 2285 groups using the BLAST algorithm for computing similarity measures. SBASE 3.0 is freely available on request to the authors or by anonymous 'ftp' file transfer from < ftp.icgeb.trieste.it >. Individual records can be retrieved with the gopher server at < icgeb.trieste.it > and with a www-server at < http:@www.icgeb.trieste.it >. Automated searching of SBASE by BLAST can be carried out with the electronic mail server < sbase@icgeb.trieste.it >. Another mail server < domain@hubi.abc.hu > assigns SBASE domain homologies on the basis of SWISS-PROT searches. A comparison of pertinent search strategies is presented.

Amino Acid Sequence↗

A novel aspect of the information content of viroids.

Viroids were found to exhibit a structural periodicity characterized by repeat units of a length of 11 or 12 (potato spindle tuber viroid group and coconut cadang-cadang viroid), 60 (apple scar skin viroid) and 80 (avocado sunblotch viroid) nucleotide residues, respectively. It is suggested that structural periodicity of viroids is an indication of their protein-binding ability.

Base Sequence↗

Plant small nuclear RNAs. V. U4 RNA is present in broad bean plants in the form of sequence variants and is base-paired with U6 RNA.

U4 RNA, which is known to play an indispensable role in pre-mRNA splicing, is present in plant nuclei, has a canonical m3 2,2,7 G cap at its 5' end and is associated with U6 RNA in snRNP particles. It occurs in broad bean in the form of a number of sequence variants. Two of these were sequenced: U4A RNA is 154 and U4B RNA is 152 nucleotides long. Sequence similarity of broad bean U4B RNA is 94 per cent to broad bean U4A RNA, 65 per cent to rat U4A RNA, 61 per cent to Drosophila U4A RNA and 50 per cent to snR14, the U4 RNA equivalent of the yeast Saccharomyces cerevisiae. Sequence conservation is much more pronounced in the 5' half of the molecule than in its 3' half. The secondary structure of both variants of broad bean U4 RNA perfectly fits with that of all other U4 RNAs sequenced so far. Nucleotide changes between broad bean U4A and U4B RNAs are restricted to molecular regions that affect the thermodynamic stability of these molecules. A model is proposed for the base pairing interaction of broad bean U4 RNA with broad bean U6 RNA. This is the first report on the structure of a plant U4 RNA.

Animals↗