PubMed Health⌕ Search

Biomedical subjects

William R Pearson

Publications and source records attributed to William R Pearson.

12 recordsLinked to original sources

The limits of protein sequence comparison?

Modern sequence alignment algorithms are used routinely to identify homologous proteins, proteins that share a common ancestor. Homologous proteins always share similar structures and often have similar functions. Over the past 20 years, sequence comparison has become both more sensitive, largely because of profile-based methods, and more reliable, because of more accurate statistical estimates. As sequence and structure databases become larger, and comparison methods become more powerful, reliable statistical estimates will become even more important for distinguishing similarities that are due to homology from those that are due to analogy (convergence). The newest sequence alignment methods are more sensitive than older methods, but more accurate statistical estimates are needed for their full power to be realized.

Algorithms↗

Nomenclature for mammalian soluble glutathione transferases.

The nomenclature for human soluble glutathione transferases (GSTs) is extended to include new members of the GST superfamily that have been discovered, sequenced, and shown to be expressed. The GST nomenclature is based on primary structure similarities and the division of GSTs into classes of more closely related sequences. The classes are designated by the names of the Greek letters: Alpha, Mu, Pi, etc., abbreviated in Roman capitals: A, M, P, and so on. (The Greek characters should not be used.) Class members are distinguished by Arabic numerals and the native dimeric protein structures are named according to their subunit composition (e.g., GST A1-2 is the enzyme composed of subunits 1 and 2 in the Alpha class). Soluble GSTs from other mammalian species can be classified in the same manner as the human enzymes, and this chapter presents the application of the nomenclature to the rat and mouse GSTs.

Animals↗

Phylogenies of glutathione transferase families.

The best known glutathione transferase family, with its class-alpha, -mu, -pi, -omega, -sigma, -theta, and -zeta subdivisions, is only one of four, or perhaps five, ancient protein families that conjugate glutathione or use a glutathione intermediate: (1) the cytoplasmic family, (2) the mitochondrial (kappa) family, (3) the microsomal (MAPEG) family, which may actually be two separate families, and (4) the fosphomycin/glyoxalase family. Although the cytoplasmic family is perhaps the most diverse, all four of these families have homologs in both prokaryotes and eukaryotes; it is striking that at least three, and perhaps as many as five, different protein folds capable of binding and positioning glutathione for a nucleophilic attack emerged more than 2 billion years ago. This chapter presents phylogenies for the four (or five) glutathione transferase families, focusing on the statistical evidence for homology (and non-homology).

Animals↗

The genome of Cryptosporidium hominis.

Cryptosporidium species cause acute gastroenteritis and diarrhoea worldwide. They are members of the Apicomplexa--protozoan pathogens that invade host cells by using a specialized apical complex and are usually transmitted by an invertebrate vector or intermediate host. In contrast to other Apicomplexans, Cryptosporidium is transmitted by ingestion of oocysts and completes its life cycle in a single host. No therapy is available, and control focuses on eliminating oocysts in water supplies. Two species, C. hominis and C. parvum, which differ in host range, genotype and pathogenicity, are most relevant to humans. C. hominis is restricted to humans, whereas C. parvum also infects other mammals. Here we describe the eight-chromosome approximately 9.2-million-base genome of C. hominis. The complement of C. hominis protein-coding genes shows a striking concordance with the requirements imposed by the environmental niches the parasite inhabits. Energy metabolism is largely from glycolysis. Both aerobic and anaerobic metabolisms are available, the former requiring an alternative electron transport system in a simplified mitochondrion. Biosynthesis capabilities are limited, explaining an extensive array of transporters. Evidence of an apicoplast is absent, but genes associated with apical complex organelles are present. C. hominis and C. parvum exhibit very similar gene complements, and phenotypic differences between these parasites must be due to subtle sequence divergence.

Animals↗

Visualization of near-optimal sequence alignments.

MOTIVATION: Mathematically optimal alignments do not always properly align active site residues or well-recognized structural elements. Most near-optimal sequence alignment algorithms display alternative alignment paths, rather than the conventional residue-by-residue pairwise alignment. Typically, these methods do not provide mechanisms for finding effectively the most biologically meaningful alignment in the potentially large set of options. RESULTS: We have developed Web-based software that displays near optimal or alternative alignments of two protein or DNA sequences as a continuous moving picture. A WWW interface to a C++ program generates near optimal alignments, which are sent to a Java Applet, which displays them in a series of alignment frames. The Applet aligns residues so that consistently aligned regions remain at a fixed position on the display, while variable regions move. The display can be stopped to examine alignment details.

Algorithms↗

Sensitivity and selectivity in protein structure comparison.

Seven protein structure comparison methods and two sequence comparison programs were evaluated on their ability to detect either protein homologs or domains with the same topology (fold) as defined by the CATH structure database. The structure alignment programs Dali, Structal, Combinatorial Extension (CE), VAST, and Matras were tested along with SGM and PRIDE, which calculate a structural distance between two domains without aligning them. We also tested two sequence alignment programs, SSEARCH and PSI-BLAST. Depending upon the level of selectivity and error model, structure alignment programs can detect roughly twice as many homologous domains in CATH as sequence alignment programs. Dali finds the most homologs, 321-533 of 1120 possible true positives (28.7%-45.7%), at an error rate of 0.1 errors per query (EPQ), whereas PSI-BLAST finds 365 true positives (32.6%), regardless of the error model. At an EPQ of 1.0, Dali finds 42%-70% of possible homologs, whereas Matras finds 49%-57%; PSI-BLAST finds 36.9%. However, Dali achieves >84% coverage before the first error for half of the families tested. Dali and PSI-BLAST find 9.2% and 5.2%, respectively, of the 7056 possible topology pairs at an EPQ of 0.1 and 19.5, and 5.9% at an EPQ of 1.0. Most statistical significance estimates reported by the structural alignment programs overestimate the significance of an alignment by orders of magnitude when compared with the actual distribution of errors. These results help quantify the statistical distinction between analogous and homologous structures, and provide a benchmark for structure comparison statistics.

Computational Biology↗

CRP: Cleavage of Radiolabeled Phosphoproteins.

The CRP (Cleavage of Radiolabeled Phosphoproteins) program guides the design and interpretation of experiments to identify protein phosphorylation sites by Edman sequencing of unseparated peptides. Traditionally, phosphorylation sites are determined by cleaving the phosphoprotein and separating the peptides for Edman 32P-phosphate release sequencing. CRP analysis of a phosphoprotein's sequence accelerates this process by omitting the separation step: given a protein sequence of interest, the CRP program performs an in silico proteolytic cleavage of the sequence and reports the predicted Edman cycles in which radioactivity would be observed if a given serine, threonine or tyrosine were phosphorylated. Experimentally observed cycles containing 32P can be compared with CRP predictions to confirm candidate sites and/or explore the ability of additional cleavage experiments to resolve remaining ambiguities. To reduce ambiguity, the phosphorylated residue (P-Tyr, P-Ser or P-Thr) can be determined experimentally, and CRP will ignore sites with alternative residues. CRP also provides simple predictions of likely phosphorylation sites using known kinase recognition motifs. The CRP interface is available at http://fasta.bioch.virginia.edu/crp.

Humans↗

Identification of residues in glutathione transferase capable of driving functional diversification in evolution. A novel approach to protein redesign.

Evolution of protein function can be driven by positive selection of advantageous nonsynonymous codon mutations that arise following gene duplication. By observing the presence and degree of site-specific positive selection for change between divergent paralogs, residue positions responsible for functional changes can be identified. We applied this analysis to genes encoding Mu class glutathione transferases, which differ widely in substrate specificities. Approximately 3% of the amino acid residue positions, both near to and distant from the active site, are under statistically significant positive selection for change. Relevant human glutathione transferase (GST) M1-1 and GST M2-2 codons were mutated. A chemically conservative threonine to serine mutation in GST M2-2 elicited a 1,000-fold increase in specific activity with the GST M1-1-specific substrate trans-stilbene oxide and a 30-fold increase with the alternative epoxide substrates styrene oxide and nitrophenyl glycidol. The reverse mutation in GST M1-1 resulted in reciprocal decreases in activity. Thus, identification of hypervariable codon positions can be a powerful aid in the redesign of protein function, lessening the requirement for extensive mutagenesis or structural knowledge and sometimes suggesting mutations that would otherwise be considered functionally conservative.

Evolution, Molecular↗

Identification and characterization of GSTT3, a third murine Theta class glutathione transferase.

A novel Theta class glutathione transferase (GST) isoenzyme from mouse termed mGSTT3 has been identified by analysis of the expressed sequence tag database. The gene encoding mGSTT3 is clustered with the mGSTT1 and mGSTT2 genes on chromosome 10 and has an exon/intron structure that is similar to that of the other Theta class genes. mGSTT3 is expressed strongly in the liver and to a decreasing extent in the kidney and testis. Recombinant mGSTT3-3 expressed in Escherichia coli had a substrate-specificity profile that differed significantly from that of GSTT1-1 and GSTT2-2 isoenzymes. A molecular model of mGSTT3 suggested that, in comparison with GSTT2, a decrease in volume of the hydrophobic substrate-binding site and the loss of the sulphate-binding pocket prevents its use of the GSTT2 substrate 1-menaphthyl sulphate.

Amino Acid Sequence↗

Getting more from less: algorithms for rapid protein identification with multiple short peptide sequences.

We describe two novel sequence similarity search algorithms, FASTS and FASTF, that use multiple short peptide sequences to identify homologous sequences in protein or DNA databases. FASTS searches with peptide sequences of unknown order, as obtained by mass spectrometry-based sequencing, evaluating all possible arrangements of the peptides. FASTF searches with mixed peptide sequences, as generated by Edman sequencing of unseparated mixtures of peptides. FASTF deconvolutes the mixture, using a greedy heuristic that allows rapid identification of high scoring alignments while reducing the total number of explored alternatives. Both algorithms use the heuristic FASTA comparison strategy to accelerate the search but use alignment probability, rather than similarity score, as the criterion for alignment optimality. Statistical estimates are calculated using an empirical correction to a theoretical probability. These calculated estimates were accurate within a factor of 10 for FASTS and 1000 for FASTF on our test dataset. FASTS requires only 15-20 total residues in three or four peptides to robustly identify homologues sharing 50% or greater protein sequence identity. FASTF requires about 25% more sequence data than FASTS for equivalent sensitivity, but additional sequence data are usually available from mixed Edman experiments. Thus, both algorithms can identify homologues that diverged 100 to 500 million years ago, allowing proteomic identification from organisms whose genomes have not been sequenced.

Algorithms↗

A strategy for the rapid identification of phosphorylation sites in the phosphoproteome.

Edman phosphate ((32)P) release sequencing provides a high sensitivity means of identifying phosphorylation sites in proteins that complements mass spectrometry techniques. We have developed a bioinformatic assessment tool, the cleavage of radiolabeled protein (CRP) program, which enables experimental identification of phosphorylation sites via (32)P labeling and Edman degradation of cleaved proteins obtained at femtomole levels. By observing the Edman cycle(s) in which radioactivity is found, candidate phosphorylation sites are identified by determining which residues occur at the observed number of cycles downstream from a peptide cleavage site. In cases where more than one residue could be responsible for the observed radioactivity, additional experiments with cleavage reagents having alternative specificities may resolve the ambiguity. Given a protein sequence and a cleavage site, CRP performs these experiments in silico, identifying resolved sites based on user-supplied experimental data, as well as suggesting combinations of reagents for additional analyses. Analysis of the PhosphoBase protein sequence database suggests that CRP data from two cleavage experiments can be used to identify unambiguously 60% of known phosphorylation sites. Data from additional cleavage experiments may increase the overall coverage to 70% of known sites. By comparing theoretical data obtained from the CRP program with (32)P release data obtained from an Edman sequencer, a known phosphorylation site was identified unambiguously and correctly. In addition, our results show that in vivo phosphorylation sites can be determined routinely by differential proteolysis analysis and Edman cycling with less than 1 fmol of protein and 1000 cpm.

Amino Acid Sequence↗

Entamoeba histolytica: sequence conservation of the Gal/GalNAc lectin from clinical isolates.

The Gal/GalNAc lectin gene of Entamoeba histolytica is a major amebic virulence protein responsible for interaction with host tissues. We investigated sequence differences in the Gal/GalNAc lectin heavy subunit in three isolates from Bangladesh and one isolate from Georgia, each of which was determined to be genetically distinct by SREHP AluI digestion. Interestingly, we observed only slight genetic diversity in the lectin gene as compared with the HM1:IMSS laboratory strain, originally a clinical isolate from Mexico. Genetic conservation of the Gal/GalNAc lectin between isolates may reflect that the lectin is under strong functional selection or possibly, that E. histolytica is a clonal population. Sequence conservation of the lectin indicates that immune responses against it should be cross-protective.

Amino Acid Sequence↗