PubMed Health⌕ Search

Biomedical subjects

Daniel Fischer

Publications and source records attributed to Daniel Fischer.

At least 19 recordsLinked to original sources

A putative novel alpha/beta hydrolase ORFan family in Bacillus.

A large number of sequences in each newly sequenced genome correspond to lineage and species-specific proteins, also known as ORFans. Amongst these ORFans, a large number are sequences with unknown structures and functions. We have identified a family of sequences, annotated as hypothetical proteins, which are specific to Bacillus and have carried out a computational study aimed at characterizing this family. Fold-recognition methods predict that these sequences belong to the alpha/beta hydrolase fold. We suggest possible catalytic triads for the ORFans and propose a hypothesis regarding the possible families within the alpha/beta hydrolase superfamily to which they may belong.

Amino Acid Sequence↗

Meta-DP: domain prediction meta-server.

UNLABELLED: Meta-DP, a domain prediction meta-server provides a simple interface to predict domains in a given protein sequence using a number of domain prediction methods. The Meta-DP is a convenient resource because through accessing a single site, users automatically obtain the results of the various domain prediction methods along with a consensus prediction. The Meta-DP is currently coupled to 10 domain prediction servers and can be extended to include any number of methods. Meta-DP can thus become a centralized repository of available methods. Meta-DP was also used to evaluate the performance of 13 domain prediction methods in the context of CAFASP-DP. AVAILABILITY: The Meta-DP server is freely available at http://meta-dp.bioinformatics.buffalo.edu and the CAFASP-DP evaluation results are available at http://cafasp4.bioinformatics.buffalo.edu/dp/update.html CONTACT: hkaur@bioinformatics.buffalo.edu SUPPLEMENTARY INFORMATION: Available at http://cafasp4.bioinformatics.buffalo.edu/dp/update.html.

Algorithms↗

LiveBench-8: the large-scale, continuous assessment of automated protein structure prediction.

We present the results of the evaluation of the latest LiveBench-8 experiment. These results provide a snapshot view of the state of the art in automated protein structure prediction, just before the 2004 CAFASP-4/CASP-6 experiments begin. The last CAFASP/CASP experiments demonstrated that automated meta-predictors entail a significant advance in the field, already challenging most human expert predictors. LiveBench-8 corroborates the superior performance of meta-predictors, which are able to produce useful predictions for over one-half of the test targets. More importantly, LiveBench-8 identifies a handful of recently developed autonomous (nonmeta) servers that perform at the very top, suggesting that further progress in the individual methods has recently been obtained.

Automation↗

Structural biology sheds light on the puzzle of genomic ORFans.

Genomic ORFans are orphan open reading frames (ORFs) with no significant sequence similarity to other ORFs. ORFans comprise 20-30% of the ORFs of most completely sequenced genomes. Because nothing can be learnt about ORFans via sequence homology, the functions and evolutionary origins of ORFans remain a mystery. Furthermore, because relatively few ORFans have been experimentally characterized, it has been suggested that most ORFans are not likely to correspond to functional, expressed proteins, but rather to spurious ORFs, pseudo-genes or to rapidly evolving proteins with non-essential roles. As a snapshot view of current ORFan structural studies, we searched for ORFans among proteins whose three-dimensional structures have been recently determined. We find that functional and structural studies of ORFans are not as underemphasized as previously suggested. These recently determined structures correspond to ORFans from all Kingdoms of life, and include proteins that have previously been functionally characterized, as well as structural genomics targets of unknown function labeled as "hypothetical proteins". This suggests that many of the ORFans in the databases are likely to correspond to expressed, functional (and even essential) proteins. Furthermore, the recently determined structures include examples of the various types of ORFans, suggesting that the functions and evolutionary origins of ORFans are diverse. Although this survey sheds some light on the ORFan mystery, further experimental studies are required to gain a better understanding of the role and origins of the tens of thousands of ORFans awaiting characterization.

Evolution, Molecular↗

The PDB-Preview database: a repository of in-silico models of 'on-hold' PDB entries.

UNLABELLED: The PDB-Preview database is a dynamic web repository of in-silico predicted three-dimensional (3D) models of experimentally determined structures that are deposited into the PDB but are not yet publicly released, and are kept 'on-hold'. The PDB-Preview database is automatically generated on a weekly basis by the bioinfo.pl meta-server, which uses top-of-the-line fold-recognition methods. The PDB-Preview provides biologists with preliminary fold assignments well before the experimentally determined 3D structures are released. AVAILABILITY: http://bioinfo.pl/PDB-Preview/.

Computer Simulation↗

The ORFanage: an ORFan database.

As each newly sequenced genome contains a significant number of protein-coding ORFs that are species-, family- or lineage-specific, many interesting questions arise about the evolution and role of these ORFs and of the genomes they are part of. We refer to these poorly conserved ORFs as singleton or paralogous ORFans if they are unique to one genome, or as orthologous ORFans if they appear only in a family of closely related organisms and have no homolog in other genomes. In order to study and classify ORFans we have constructed the ORFanage, an ORFan database. This database consists of the predicted ORFs in fully sequenced microbial genomes, and enables searching for the three types of ORFans in any subset of the genomes chosen by the user. The ORFanage could help in choosing interesting targets for further genomic and evolutionary studies. The ORFanage is accessible via http://www.bioinformatics.buffalo. edu/ORFanage.

Computational Biology↗

Analysis of singleton ORFans in fully sequenced microbial genomes.

Singleton sequence ORFans are orphan ORFs (open reading frames) that have no detectable sequence similarity to any other sequence in the databases. ORFans are of particular interest not only as evolutionary puzzles but also because we can learn little about them using bioinformatics tools. Here, we present a first systematic analysis of singleton ORFans in the first 60 fully sequenced microbial genomes. We show that although ORFans have been underemphasized, the number of ORFans is steadily growing, currently accounting for 23,634 sequences. At the same time, the percentage of ORFans as a fraction of all sequences is slowly diminishing, and is currently about 14%. Short ORFans comprise about 61% of all ORFans. The abundance of short ORFans may be due to a yet unexplained artifact. The data also suggest that the number of longer ORFans may soon diminish as more genomes of closely related organisms become available. To better address the questions about the functions and origins of ORFans, we propose to focus further studies on the longer ORFans, with emphasis on three new types of ORFans: ORFan modules, paralogous ORFans, and orthologous ORFans. We conclude that the large number of ORFans reflects an intrinsic property of the genetic material not yet fully understood. Further computational and experimental studies aimed at understanding Nature's protein diversity should also include ORFans.

Genome↗

3D-Jury: a simple approach to improve protein structure predictions.

MOTIVATION: Consensus structure prediction methods (meta-predictors) have higher accuracy than individual structure prediction algorithms (their components). The goal for the development of the 3D-Jury system is to create a simple but powerful procedure for generating meta-predictions using variable sets of models obtained from diverse sources. The resulting protocol should help to improve the quality of structural annotations of novel proteins. RESULTS: The 3D-Jury system generates meta-predictions from sets of models created using variable methods. It is not necessary to know prior characteristics of the methods. The system is able to utilize immediately new components (additional prediction providers). The accuracy of the system is comparable with other well-tuned prediction servers. The algorithm resembles methods of selecting models generated using ab initio folding simulations. It is simple and offers a portable solution to improve the accuracy of other protein structure prediction protocols. AVAILABILITY: The 3D-Jury system is available via the Structure Prediction Meta Server (http://BioInfo.PL/Meta/) to the academic community. SUPPLEMENTARY INFORMATION: 3D-Jury is coupled to the continuous online server evaluation program, LiveBench (http://BioInfo.PL/LiveBench/)

Algorithms↗

3D-SHOTGUN: a novel, cooperative, fold-recognition meta-predictor.

To gain a better understanding of the biological role of proteins encoded in genome sequences, knowledge of their three-dimensional (3D) structure and function is required. The computational assignment of folds is becoming an increasingly important complement to experimental structure determination. In particular, fold-recognition methods aim to predict approximate 3D models for proteins bearing no sequence similarity to any protein of known structure. However, fully automated structure-prediction methods can currently produce reliable models for only a fraction of these sequences. Using a number of semiautomated procedures, human expert predictors are often able to produce more and better predictions than automated methods. We describe a novel, fully automatic, fold-recognition meta-predictor, named 3D-SHOTGUN, which incorporates some of the strategies human predictors have successfully applied. This new method is reminiscent of the so-called cooperative algorithms of Computer Vision. The input to 3D-SHOTGUN are the top models predicted by a number of independent fold-recognition servers. The meta-predictor consists of three steps: (i) assembly of hybrid models, (ii) confidence assignment, and (iii) selection. We have applied 3D-SHOTGUN to an unbiased test set of 77 newly released protein structures sharing no sequence similarity to proteins previously released. Forty-six correct rank-1 predictions were obtained, 30 of which had scores higher than that of the first incorrect prediction-a significant improvement over the performance of all individual servers. Furthermore, the predicted hybrid models were, on average, more similar to their corresponding native structures than those produced by the individual servers. This opens the possibility of generating more accurate, full-atom homology models for proteins with no sequence similarity to proteins of known structure. These improvements represent a step forward toward the wider applicability of fully automated structure-prediction methods at genome scales.

Algorithms↗

LiveBench-6: large-scale automated evaluation of protein structure prediction servers.

The aim of the LiveBench experiment is to provide a continuous evaluation of structure prediction servers in order to inform potential users about the current state-of-the-art structure prediction tools and in order to help the developers to analyze and improve the services. This round of the experiment was conducted in parallel to the blind CAFASP-3 evaluation experiment. The data collected almost simultaneously enables the comparison of servers on two different benchmark sets. The number of servers has doubled from the last evaluated LiveBench-4 experiment completed in April 2002, just before the beginning of CAFASP-3. This can be partially attributed to the rapid development in the area of meta-predictors (consensus servers). The current results confirm the high sensitivity and specificity of the meta-predictors. Nevertheless, the comparison between the autonomous (not meta) servers participating in the last CAFASP-2 and LiveBench-2 experiment and the current set of autonomous servers demonstrates that progress has been made also in sequence structure fitting functions. In addition to the growing number of participants, the current experiment marks the introduction of new evaluation procedures, which are aimed to correlate better with functional characteristics of models.

Amino Acids↗

3DS3 and 3DS5 3D-SHOTGUN meta-predictors in CAFASP3.

The performance of the 3DS3 and 3DS5 3D-SHOTGUN meta-predictors in CAFASP3 is reported. The 3D-SHOTGUN meta-predictors are fully automatic fold recognition servers that attempt to incorporate into the prediction process a number of successful strategies that human predictors often apply. Namely, the input to 3D-SHOTGUN are the top five models predicted by a number of independent fold recognition servers and its output are hybrid models, assembled by using the recurrent structural information from the input models. The resulting hybrid models are, on average, more accurate and more complete than the input models. When evaluated on a large set of prediction targets, the 3D-SHOTGUN servers show increased sensitivities and significantly better specificities. For CAFASP3, the 3DS3 and 3DS3 and 3DS5 used a preliminary implementation of the 3D-SHOTGUN method, which lacked a refinement step. Although this did not have a significant effect on the easier targets, for the hardest prediction targets, where the input models had significant structural conflicts, the 3D-SHOTGUN models contained a number of non-native-like features such as fragmentation and overlaps. The CAFASP3 evaluation identified the 3D-SHOTGUN meta-predictors within the top three most sensitive and most specific servers. A fully automated refinement step to the 3D-SHOTGUN method is currently being implemented, and preliminary results indicate that in addition to "cleaning up" such undesirable features, it is able to further increase the accuracy of the resulting models.

Computational Biology↗

CAFASP3: the third critical assessment of fully automated structure prediction methods.

We present the results of the fully automated CAFASP3 experiment, which was carried out in parallel with CASP5, using the same set of prediction targets. CAFASP participation is restricted to fully automatic structure prediction servers. The servers' performance is evaluated by using previously announced, objective, reproducible and fully automated evaluation methods. More than 60 servers participated in CAFASP3, covering all categories of structure prediction. As in the previous CAFASP2 experiment, it was possible to identify a group of 5-10 top performing independent servers. This group of top performing independent servers produced relatively accurate models for all the 32 "Homology Modeling" targets, and for up to 43% of the 30 "Fold Recognition" targets. One of the most important results of CAFASP3 was the realization of the value of all the independent servers as a group, as evidenced by the superior performance of "meta-predictors" (defined here as predictors that make use of the output of other CAFASP servers). The performance of the best automated meta-predictors was roughly 30% higher than that of the best independent server. More significantly, the performance of the best automated meta-predictors was comparable with that of the best 5-10 human CASP predictors. This result shows that significant progress has been achieved in automatic structure prediction and has important implications to the prospects of automated structure modeling in the context of structural genomics.

Computational Biology↗

Modeling three-dimensional protein structures for CASP5 using the 3D-SHOTGUN meta-predictors.

Full-atom models were generated for all CASP5 targets by using the fully automated 3D-SHOTGUN fold recognition meta-predictors (Fischer D, Proteins 2003;51:434-441). The 3D-SHOTGUN meta-predictors assemble hybrid 3D models by combining structural information of a number of independently generated, fold recognition models. At the time CASP5 took place, the 3D-SHOTGUN servers generated unrefined C(alpha)-only models. Fischer's participation in CASP had three main goals. The first was to test the value of using 3D-SHOTGUN models as input to a refinement procedure. The second goal was to test whether human intervention could result in a better performance than that of the automated servers. The third goal was to evaluate which human procedures, not yet implemented within the 3D-SHOTGUN servers, can be implemented in the future. For CASP5, our group's predictions applied a very simple approach using the multiple parent option of the Modeller program (Sali and Blundell, J Mol Biol 1993;234:779-815). The input to Modeller was different combinations of the unrefined 3D-SHOTGUN models and the sequence-template alignments used by 3D-SHOTGUN's assembly step. Our evaluation of the accuracies of the refined versus the SHOTGUN models shows that the refined models were consistently slightly more accurate than SHOTGUN's. For a few targets, the manual use of the information from the CAFASP servers resulted in better human models. This manual intervention was particularly valuable in the identification of domains, still a difficult feature for automated servers. The CASP5 results indicate that 3D-SHOTGUN's hybrid models can be a valuable starting point for full-atom refinement and that the resulting refined models are, on average, more accurate than those produced by the servers. Thus, we conclude that our three goals were achieved. A preliminary automated version of the refinement procedure, named SHGUM, is now available.

Computational Biology↗

Twenty thousand ORFan microbial protein families for the biologist?

The genomes of most newly sequenced organisms contain a significant fraction of ORFs (open reading frames) that match no other sequence in the databases. We refer to these singleton ORFs as sequence ORFans. Because little can be learned about ORFans by homology, the origin and functions of ORFans remain a mystery. However, in this era of full genome sequencing, it seems that ORFans have been underemphasized. In this minireview, we draw attention to the increasing number of ORFans and to the consequences of this growth to biological research in the postgenomic era.

Animals↗

The 2002 Olympic Games of protein structure prediction.

The summer of every even year is considered by the protein structure prediction community as the Olympic Games season, because in addition to a number of continuous benchmarking experiments such as LiveBench, much effort is invested in the blind prediction experiments CASP and CAFASP. Here we report the major advances registered in the field since the last Games of 2000, as measured by the recently completed LiveBench-4 experiment. These results provide a timely measure of the capabilities of current methods and of their expected performance in the upcoming CASP-5 and CAFASP-3 experiments. We also describe the initiation of the two new, community-wide experiments, PDB-CAFASP and MR-CAFASP. These new experiments extend the scope of previous efforts and may have important implications for structural genomics.

Computational Biology↗