PubMed Health⌕ Search

Biomedical subjects

T Hubbard

Publications and source records attributed to T Hubbard.

At least 19 recordsLinked to original sources

Ensembl 2002: accommodating comparative genomics.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organise biology around the sequences of large genomes. It is a comprehensive source of stable automatic annotation of human, mouse and other genome sequences, available as either an interactive web site or as flat files. Ensembl also integrates manually annotated gene structures from external sources where available. As well as being one of the leading sources of genome annotation, Ensembl is an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements. These range from sequence analysis to data storage and visualisation and installations exist around the world in both companies and at academic sites. With both human and mouse genome sequences available and more vertebrate sequences to follow, many of the recent developments in Ensembl have focusing on developing automatic comparative genome analysis and visualisation.

Animals↗

The Ensembl genome database project.

The Ensembl (http://www.ensembl.org/) database project provides a bioinformatics framework to organise biology around the sequences of large genomes. It is a comprehensive source of stable automatic annotation of the human genome sequence, with confirmed gene predictions that have been integrated with external data sources, and is available as either an interactive web site or as flat files. It is also an open source software engineering project to develop a portable system able to handle very large genomes and associated requirements from sequence analysis to data storage and visualisation. The Ensembl site is one of the leading sources of human genome sequence annotation and provided much of the analysis for publication by the international human genome project of the draft genome. The Ensembl system is being installed around the world in both companies and academic sites on machines ranging from supercomputers to laptops.

Computational Biology↗

Initial sequencing and analysis of the human genome.

The human genome holds an extraordinary trove of information about human development, physiology, medicine and evolution. Here we report the results of an international collaboration to produce and make freely available a draft sequence of the human genome. We also present an initial analysis of the data, describing some of the insights that can be gleaned from the sequence.

Animals↗

The Interdisciplinary Generalist Curriculum Project at Eastern Virginia Medical School.

The proposed Interdisciplinary Generalist Curriculum (IGC) Project at Eastern Virginia Medical School (hereafter Eastern Virginia) intended to encourage students to select generalist disciplines by featuring generalist role models, focusing on patients' perspectives, teaching generalist skills early, providing care to indigent and other populations, and emphasizing students' personal and professional development. To do so, Eastern Virginia proposed that collaborative interdisciplinary groups of faculty plan and oversee the implementation of first- and second-year students' early clinical experiences in generalists' offices as integrated with new and revised first- and second-year courses, the coordination of generalist curricula longitudinally from year one through year four, and the provision of appropriate faculty development. With minor exceptions described, the project was implemented as proposed. The project did have desirable effects, both intended and unexpected. The curricular changes made in the project will remain.

Curriculum↗

Analysis and assessment of ab initio three-dimensional prediction, secondary structure, and contacts prediction.

CASP3 saw a substantial increase in the volume of ab initio 3D prediction data, with 507 datasets for fifteen selected targets and sixty-one groups participating. As with CASP2, methods ranged from computationally intensive strategies that attempt to recreate the physical and chemical forces involved in protein folding to the more recent knowledge-based approaches. These exploit information from the structure databases, extracting potentially similar fragments and/or distance constraints derived from multiple sequence alignments. The knowledge-based approaches generally gave more consistently successful predictions across the range of targets, particularly that of the Baker group (Bystroff and Baker, J Mol Biol 1998;281:565-577; Simons et al. Proteins Suppl 1999;3:171-176), which used a fragment library. In the secondary structure prediction category, the most successful approaches built on the concepts used in PHD (Rost et al. Comput Appl Biosci 1994;10:53-60), an accepted standard in this field. Like PHD, they exploit neural networks but have different strategies for incorporating multiple sequence data or position-dependent weight matrices for training the networks. Analysis of the contact data, for which only six groups participated, suggested that as yet this data provides a rather weak signal. However, in combination with other types of prediction data it can sometimes be a useful constraint for identifying the correct structure.

Animals↗

Sequence comparisons using multiple sequences detect three times as many remote homologues as pairwise methods.

The sequences of related proteins can diverge beyond the point where their relationship can be recognised by pairwise sequence comparisons. In attempts to overcome this limitation, methods have been developed that use as a query, not a single sequence, but sets of related sequences or a representation of the characteristics shared by related sequences. Here we describe an assessment of three of these methods: the SAM-T98 implementation of a hidden Markov model procedure; PSI-BLAST; and the intermediate sequence search (ISS) procedure. We determined the extent to which these procedures can detect evolutionary relationships between the members of the sequence database PDBD40-J. This database, derived from the structural classification of proteins (SCOP), contains the sequences of proteins of known structure whose sequence identities with each other are 40% or less. The evolutionary relationships that exist between those that have low sequence identities were found by the examination of their structural details and, in many cases, their functional features. For nine false positive predictions out of a possible 432,680, i.e. at a false positive rate of about 1/50,000, SAM-T98 found 35% of the true homologous relationships in PDBD40-J, whilst PSI-BLAST found 30% and ISS found 25%. Overall, this is about twice the number of PDBD40-J relations that can be detected by the pairwise comparison procedures FASTA (17%) and GAP-BLAST (15%). For distantly related sequences in PDBD40-J, those pairs whose sequence identity is less than 30%, SAM-T98 and PSI-BLAST detect three times the number of relationships found by the pairwise methods.

Databases, Factual↗

Using neural networks for prediction of the subcellular location of proteins.

Neural networks have been trained to predict the subcellular location of proteins in prokaryotic or eukaryotic cells from their amino acid composition. For three possible subcellular locations in prokaryotic organisms a prediction accuracy of 81% can be achieved. Assigning a reliability index, 33% of the predictions can be made with an accuracy of 91%. For eukaryotic proteins (excluding plant sequences) an overall prediction accuracy of 66% for four locations was achieved, with 33% of the sequences being predicted with an accuracy of 82% or better. With the subcellular location restricting a protein's possible function, this method should be a useful tool for the systematic analysis of genome data and is available via a server on the world wide web.

Amino Acids↗

GLASS: a tool to visualize protein structure prediction data in three dimensions and evaluate their consistency.

When a protein sequence does not share any significant sequence similarity with a protein of known structure, homology modeling cannot be applied. However, many novel and interesting methods, such as secondary structure prediction, fold recognition, and prediction of long-range interactions, are being developed and have been shown to be reasonably successful in predicting protein structures from sequence data and evolutionary information. The a priori evaluation of the correctness of a prediction obtained by one of these methods is however often problematic. Consequently, it is important to use all available information provided by as many different methods as possible and all the available experimental data about the protein of interest, since the consistency of the results is indicative of the reliability of the prediction. Hence the need has arisen for suitable tools able to compare results provided by different methods and evaluate their consistency. We have therefore constructed GLASS, a general platform to read, visualize, compare, and evaluate prediction results from many different sources and to project these prediction results into three dimensions. In addition, GLASS allows the comparison of selected parameters calculated for a model with the distribution observed in real protein structures, thus providing an easy way to test new methods for evaluating the likelihood of different structural models. GLASS can be considered as a "workbench" for structural predictions useful to both experimentalists and theoreticians.

Amino Acid Sequence↗

SPEM: a parser for EMBL style flat file database entries.

SUMMARY: We present a set of Perl modules for the flexible and robust parsing and editing of EMBL/SWISS-PROT databases. AVAILABILITY: The Web page at http://www.sanger.ac. uk/Software/PerlModule/ provides information about downloading the SPEM and PrEMBL modules, and provides links to documentation and example code.

Databases, Factual↗

Intermediate sequences increase the detection of homology between sequences.

Two homologous sequences, which have diverged beyond the point where their homology can be recognised by a simple direct comparison, can be related through a third sequence that is suitably intermediate between the two. High scores, for a sequence match between the first and third sequences and between the second and the third sequences, imply that the first and second sequences are related even though their own match score is low. We have tested the usefulness of this idea using a database that contains the sequences of 971 protein domains whose structures are known and whose residue identities with each other are some 40% or less (PDB40D). On the basis of sequence and structural information, 2143 pairs of these sequences are known to have an evolutionary relationship. FASTA, in an all-against-all comparison of the sequences in the database, detected 320 (15%) of these relationships as well as three false positive (i.e. 1% error rate). Using intermediate sequences found by FASTA matches of PDB40D sequences to those in the large non-redundant OWL database we could detect 550 evolutionary relationships with an error rate of 1%. This means the intermediate sequence procedure increases the ability to recognise the evolutionary relationships amongst the PDB40D sequences by 70%.

Amino Acid Sequence↗

Protein folds in the all-beta and all-alpha classes.

Analysis of the structures in the Protein Databank, released in June 1996, shows that the number of different protein folds, i.e. the number of different arrangements of major secondary structures and/or chain topologies, is 327. Of these folds, approximately 25% belong to the all-alpha class, 20% belong to the all-beta class, 30% belong to the alpha/beta class, and 25% belong to the alpha + beta class. We describe the types of folds now known for the all-beta and all-alpha classes, emphasizing those that have been discovered recently. Detailed theories for the physical determinants of the structures of most of these folds now exist, and these are reviewed.

Databases, Factual↗

Update on protein structure prediction: results of the 1995 IRBM workshop.

Computational tools for protein structure prediction are of great interest to molecular, structural and theoretical biologists due to a rapidly increasing number of protein sequences with no known structure. In October 1995, a workshop was held at IRBM to predict as much as possible about a number of proteins of biological interest using ab initio prediction of fold recognition methods. 112 protein sequences were collected via an open invitation for target submissions. 17 were selected for prediction during the workshop and for 11 of these a prediction of some reliability could be made. We believe that this was a worthwhile experiment showing that the use of a range of independent prediction methods and thorough use of existing databases can lead to credible and useful ab initio structure predictions.

Animals↗

Discovery of an orally bioavailable NK1 receptor antagonist, (2S,3S)-(2-methoxy-5-tetrazol-1-ylbenzyl)(2-phenylpiperidin-3-yl)amine (GR203040), with potent antiemetic activity.

The antiemetic, pharmacokinetic, and metabolic profile of CP-99,994, a potent NK1 receptor antagonist, has been carefully evaluated. As a result we began a medicinal chemistry program which initially identified a 3-furanyl analogue (6) with improved antiemetic potency and a methyl sulfone (5) with enhanced metabolic stability and oral bioavailability. The improved pharmacokinetic profile of methyl sulfone (5) was associated with its low lipophilicity, and a therefore a number of heterocyclic analogues with reduced log D were synthesized. Out of this program emerged 19 (GR203040), a tetrazolyl-substituted analogue. Tetrazole 19 inhibits radiation-induced emesis in the ferret with high potency when administered both subcutaneously and orally, has a long duration of action, and has high oral bioavailability in the dog. Tetrazole 19 is currently undergoing evaluation as a novel approach for the control of emesis associated with, for example, cancer chemotherapy.

Animals↗