PubMed Health⌕ Search

Biomedical subjects

Cyrus Chothia

Publications and source records attributed to Cyrus Chothia.

13 recordsLinked to original sources

VH gene segments in the mouse and human genomes.

We have examined the mouse genome sequence to determine its VH gene segment repertoire. In all, 141 segments are mapped to a 3 Mb region of chromosome 12. There is evidence that 92 of these are functional in the mouse strain used for the genome sequence, C57BL/6J; 12 are functional in other mouse strains, and 37 are pseudogenes. The mouse VH gene segment repertoire is therefore twice the size of that in humans. The mouse and human loci bear no large-scale similarity to each other. The 104 functional segments belong to one of the 15 known sequence subgroups, which have been further clustered into eight sets here. Seven of these sets, comprising 101 sequences, are related to five of the human VH families and have the same canonical structures in their hypervariable regions. Duplication of members of one set in the distal half of the locus is mainly responsible for the larger size of the mouse repertoire. Phylogenetic analysis of the VH segments indicates that most of the sequences in the human and mouse VH loci have arisen subsequent to the divergence of the two organisms from their common ancestor.

Amino Acid Sequence↗

SCOP database in 2004: refinements integrate structure and sequence family data.

The Structural Classification of Proteins (SCOP) database is a comprehensive ordering of all proteins of known structure, according to their evolutionary and structural relationships. Protein domains in SCOP are hierarchically classified into families, superfamilies, folds and classes. The continual accumulation of sequence and structural data allows more rigorous analysis and provides important information for understanding the protein world and its evolutionary repertoire. SCOP participates in a project that aims to rationalize and integrate the data on proteins held in several sequence and structure databases. As part of this project, starting with release 1.63, we have initiated a refinement of the SCOP classification, which introduces a number of changes mostly at the levels below superfamily. The pending SCOP reclassification will be carried out gradually through a number of future releases. In addition to the expanded set of static links to external resources, available at the level of domain entries, we have started modernization of the interface capabilities of SCOP allowing more dynamic links with other databases. SCOP can be accessed at http://scop.mrc-lmb.cam.ac.uk/scop.

Animals↗

The SUPERFAMILY database in 2004: additions and improvements.

The SUPERFAMILY database provides structural assignments to protein sequences and a framework for analysis of the results. At the core of the database is a library of profile Hidden Markov Models that represent all proteins of known structure. The library is based on the SCOP classification of proteins: each model corresponds to a SCOP domain and aims to represent an entire superfamily. We have applied the library to predicted proteins from all completely sequenced genomes (currently 154), the Swiss-Prot and TrEMBL databases and other sequence collections. Close to 60% of all proteins have at least one match, and one half of all residues are covered by assignments. All models and full results are available for download and online browsing at http://supfam.org. Users can study the distribution of their superfamily of interest across all completely sequenced genomes, investigate with which other superfamilies it combines and retrieve proteins in which it occurs. Alternatively, concentrating on a particular genome as a whole, it is possible first, to find out its superfamily composition, and secondly, to compare it with that of other genomes to detect superfamilies that are over- or under-represented. In addition, the webserver provides the following standard services: sequence search; keyword search for genomes, superfamilies and sequence identifiers; and multiple alignment of genomic, PDB and custom sequences.

Animals↗

Structure, function and evolution of multidomain proteins.

Proteins are composed of evolutionary units called domains; the majority of proteins consist of at least two domains. These domains and nature of their interactions determine the function of the protein. The roles that combinations of domains play in the formation of the protein repertoire have been found by analysis of domain assignments to genome sequences. Additional findings on the geometry of domains have been gained from examination of three-dimensional protein structures. Future work will require a domain-centric functional classification scheme and efforts to determine structures of domain combinations.

Computer Simulation↗

The linked conservation of structure and function in a family of high diversity: the monomeric cupredoxins.

The monomeric cupredoxins are a highly divergent family of copper binding electron transport proteins that function in photosynthesis and respiration. To determine how function and structure are conserved in the context of large sequence differences, we have carried out a detailed analysis of the cupredoxins of known structure and their sequence homologs. The common structure of the cupredoxins is formed by a sandwich of two beta sheets which support a copper binding site. The structure of the deeply buried core is intimately coupled to the binding site on the surface of the protein; in each protein the conserved regions form one continuous substructure that extends from the surface active site and through the center of the molecule. Residues around the active site are conserved for functional reasons, while those deeper in the structure will be conserved for structural reasons. Together the two sets support each other.

Azurin↗

Exegesis: a procedure to improve gene predictions and its use to find immunoglobulin superfamily proteins in the human and mouse genomes.

Exegesis is a procedure to refine the gene predictions that are produced for complex genomes, e.g. those of humans and mice. It uses the program Genewise, sequences determined by experiment, experimental maps of gene segment libraries and a new browser that allows the user to rapidly inspect and compare multiple gene maps to regions of genomic sequences. The procedure should be of general use. Here, we use the procedure to find members of the immunoglobulin superfamily in the human and mouse genomes. To do this, Exegesis was used to process the original gene predictions from the automated Ensembl annotation pipeline. Exegesis produced (i) many more complete genes and new transcripts and (ii) a mapping of the immunoglobulin and T cell receptor gene libraries to the genome, which are largely absent in the Ensembl set.

Animals↗

Evolution of the protein repertoire.

Most proteins have been formed by gene duplication, recombination, and divergence. Proteins of known structure can be matched to about 50% of genome sequences, and these data provide a quantitative description and can suggest hypotheses about the origins of these processes.

Amino Acid Sequence↗

The immunoglobulin superfamily in Drosophila melanogaster and Caenorhabditis elegans and the evolution of complexity.

Drosophila melanogaster is an arthropod with a much more complex anatomy and physiology than the nematode Caenorhabditis elegans. We investigated one of the protein superfamilies in the two organisms that plays a major role in development and function of cell-cell communication: the immunoglobulin superfamily (IgSF). Using hidden Markov models, we identified 142 IgSF proteins in Drosophila and 80 in C. elegans. Of these, 58 and 22, respectively, have been previously identified by experiments. On the basis of homology and the structural characterisation of the proteins, we can suggest probable types of function for most of the novel proteins. Though overall Drosophila has fewer genes than C. elegans, it has many more IgSF cell-surface and secreted proteins. Half the IgSF proteins in C. elegans and three quarters of those in Drosophila have evolved subsequent to the divergence of the two organisms. These results suggest that the expansion of this protein superfamily is one of the factors that have contributed to the formation of the more complex physiological features that are found in Drosophila.

Animals↗

Sequence conservation in families whose members have little or no sequence similarity: the four-helical cytokines and cytochromes.

Proteins for which there are good structural, functional and genetic similarities that imply a common evolutionary origin, can have sequences whose similarities are low or undetectable by conventional sequence comparison procedures. Do these proteins have sequence conservation beyond the simple conservation of hydrophobic and hydrophilic character at specific sites and if they do what is its nature? To answer these questions we have analysed the structures and sequences of two superfamilies: the four-helical cytokines and cytochromes c'-b(562). Members of these superfamilies have sequence similarities that are either very low or not detectable. The cytokine superfamily has within it a long chain family and a short chain family. The sequences of known representative structures of the two families were aligned using structural information. From these alignments we identified the regions that conserve the same main-chain conformation: the common core (CC). For members of the same family, the CC comprises some 50% of the individual structures; for the combination of both families it is 30%. We added homologous sequences to the structural alignment. Analysis of the residues occurring at sites within the CCs showed that 30% have little or no conservation, whereas about 40% conserve the polar/neutral or hydrophobic/neutral character of their residues. The remaining 30% conserve hydrophobic residues with strong or medium limitations on their volume variations. Almost all of these residues are found at sites that form the "buried spine" of each helix (at sites i, i+3, i+7, i+10, etc., or i, i+4, i+7, i+11, etc.) and they pack together at the centre of each structure to give a pattern of residue-residue contacts that is almost absolutely conserved. These CC conserved hydrophobic residues form only 10-15% of all the residues in the individual structures.A similar analysis of the cytochromes c'-b(562), which bind haem and have a very different function to that of the cytokines, gave very similar results. Again some 30% of the CC residues have hydrophobic residues with strong or medium conservation. Most of these form the buried spine of each helix and play the same role as those in the cytokines. The others, and some spine residues bind the haem co-factor.

Automation↗

The geometry of domain combination in proteins.

Most proteins in genomes are the result of the recombination of two or more domains. It has been found that if proteins are formed by a combination of domains from superfamilies A and B, then the domains may occur in the sequential order AB or BA but only in about 2% of cases do they occur in both sequential orders. The classical Rossmann domains of known structure are combined with catalytic domains from seven different superfamilies. In addition, there are eight cases where structures with both AB and BA domain combinations are known. For these two sets of structures, we analysed: (i) the relative orientation of the domains; (ii) the type of domain connection; (iii) the structure of the interdomain links; and (iv) domain function. The results of this analysis indicate that in most cases domain order is conserved because recombination of the domains has only occurred once during the course of evolution. Functional reasons become important when the domain connections are short. In seven out of the eight known cases where domains are combined in the AB and BA sequential orders they have different geometrical relationships that give them different functional properties.

Animals↗

SCOP database in 2002: refinements accommodate structural genomics.

The SCOP (Structural Classification of Proteins) database is a comprehensive ordering of all proteins of known structure, according to their evolutionary and structural relationships. Protein domains in SCOP are grouped into species and hierarchically classified into families, superfamilies, folds and classes. Recently, we introduced a new set of features with the aim of standardizing access to the database, and providing a solid basis to manage the increasing number of experimental structures expected from structural genomics projects. These features include: a new set of identifiers, which uniquely identify each entry in the hierarchy; a compact representation of protein domain classification; a new set of parseable files, which fully describe all domains in SCOP and the hierarchy itself. These new features are reflected in the ASTRAL compendium. The SCOP search engine has also been updated, and a set of links to external resources added at the level of domain entries. SCOP can be accessed at http://scop.mrc-lmb.cam.ac.uk/scop.

Animals↗

SUPERFAMILY: HMMs representing all proteins of known structure. SCOP sequence searches, alignments and genome assignments.

The SUPERFAMILY database contains a library of hidden Markov models representing all proteins of known structure. The database is based on the SCOP 'superfamily' level of protein domain classification which groups together the most distantly related proteins which have a common evolutionary ancestor. There is a public server at http://supfam.org which provides three services: sequence searching, multiple alignments to sequences of known structure, and structural assignments to all complete genomes. Given an amino acid or nucleotide query sequence the server will return the domain architecture and SCOP classification. The server produces alignments of the query sequences with sequences of known structure, and includes multiple alignments of genome and PDB sequences. The structural assignments are carried out on all complete genomes (currently 59) covering approximately half of the soluble protein domains. The assignments, superfamily breakdown and statistics on them are available from the server. The database is currently used by this group and others for genome annotation, structural genomics, gene prediction and domain-based genomic studies.

Amino Acid Sequence↗

Comparison of the small molecule metabolic enzymes of Escherichia coli and Saccharomyces cerevisiae.

The comparison of the small molecule metabolism pathways in Escherichia coli and Saccharomyces cerevisiae (yeast) shows that 271 enzymes are common to both organisms. These common enzymes involve 384 gene products in E. coli and 390 in yeast, which are between one half and two thirds of the gene products of small molecule metabolism in E. coli and yeast, respectively. The arrangement and family membership of the domains that form all or part of 374 E. coli sequences and 343 yeast sequences was determined. Of these, 70% consist entirely of homologous domains, and 20% have homologous domains linked to other domains that are unique to E. coli, yeast, or both. Over two thirds of the enzymes common to the two organisms have sequence identities between 30% and 50%. The remaining groups include 13 clear cases of nonorthologous displacement. Our calculations show that at most one half to two thirds of the gene products involved in small molecule metabolism are common to E. coli and yeast. We have shown that the common core of 271 enzymes has been largely conserved since the separation of prokaryotes and eukaryotes, including modifications for regulatory purposes, such as gene fusion and changes in the number of isozymes in one of the two organisms. Only one fifth of the common enzymes have nonhomologous domains between the two organisms. Around the common core very different extensions have been made to small molecule metabolism in the two organisms.

Databases, Genetic↗