PubMed HealthSearch

Biomedical subjects

D J States

Publications and source records attributed to D J States.

10 recordsLinked to original sources

Molecular sequence accuracy: analysing imperfect data.

Molecular sequences are experimentally derived data that can be expected to contain errors as a result of diverse phenomena such as biological variation, molecular cloning artifacts, imperfect sequence determination, and data handling during contig assembly. Errors will affect the reliability of database searches and sequence alignments, but their impact may be minimized by the use of analytical techniques that anticipate that the data will be imperfect.

Amino Acid Sequence

Molecular sequence accuracy and the analysis of protein coding regions.

Molecular sequences, like all experimental data, have finite error rates. The impact of errors on the information content of molecular sequence data is dependent on the analytic paradigm used to interpret the data. We studied the impact of nucleic acid sequence errors on the ability to align predicted amino acid sequences with the sequences of related proteins. We found that with a simultaneous translation and alignment algorithm, identification of sequence homologies is resilient to the introduction of random errors. Proteins with greater than 30% sequence identity can be reliably recognized even in the presence of 1% frameshifting (insertion or deletion) error rates and 5% base substitution rates. Incorporation of prior knowledge about the location and characteristics of errors improves tolerance to error of amino acid sequence alignments. Similarly, inclusion of prior knowledge of biased codon utilization by yeast (Saccharomyces cerevisiae) allows reliable detection of correct reading frames in yeast sequences even in the presence of 5% substitution and 1% frameshift errors.

Algorithms

Structure of the human neutrophil elastase gene.

The gene for human neutrophil elastase (NE), a powerful serine protease carried by blood neutrophils and capable of destroying most connective tissue proteins, was cloned from a genomic DNA library of a normal individual. The NE gene consists of 5 exons and 4 introns included in a single copy 4-kilobase segment of chromosome 11 at q14. The coding exons of the NE gene predict a primary translation product of 267 residues including a 29-residue N-terminal precursor peptide and a 20-residue C-terminal precursor peptide. Analysis of the N-terminal peptide sequence suggests it contains a 27-residue "pre" signal peptide followed by a "proN" dipeptide, similar to that of other blood cell lysosomal proteases. The sequences for the mature 218-residue NE protein are included in exons II-V. The 5'-flanking region of the gene includes typical TATA, CAAT, and GC sequences within 61 base pairs (bp) of the cap site. The sequence 1.5 kilobases 5' to exon I contains several interesting repetitive sequences including six tandem repeats of unique 52- or 53-bp sequences. The 5'-flanking region also contains a 19-bp segment with 90% homology to a segment of the 5'-flanking region of the human myeloperoxidase (MPO) gene, a gene also expressed in bone marrow precursor cells and a protein stored in the same neutrophil granules as NE. In addition, like the MPO gene, the NE 5'-flanking region has several regions with greater than or equal to 75% homology to sequences 5' to c-myc, but there is no overlap between the NE-c-myc and MPO-c-myc homologous sequences.

Amino Acid Sequence

Electrostatic effects and hydrogen exchange behaviour in proteins. The pH dependence of exchange rates in lysozyme.

The pH dependence of the exchange rates for a number of tryptophan and amide hydrogen atoms in hen egg-white lysozyme has been determined at temperatures well below the thermal denaturation temperature. The pH behaviour of each hydrogen is unique and can differ markedly from that of simple compounds. A model for electrostatic effects in proteins is described and used to explain a number of the features of the pH dependence of the exchange rates of certain hydrogens. The results indicate that exchange takes place from a conformation of the protein closely similar to that of the native protein, with local fluctuations providing the mechanism for exchange. For the more-buried hydrogens at low pH values there is a general increase in the exchange rates caused by the decreasing stability of the protein as calculated from the electrostatic model. The analysis shows how evidence from hydrogen exchange studies can be used to provide information about electrostatic interactions in localized regions of proteins. A description of the electrostatic model and some applications are given in the Appendix.

Amino Acid Sequence

Conformations of intermediates in the folding of the pancreatic trypsin inhibitor.

Intermediates in the folding pathway of the bovine pancreatic trypsin inhibitor (PTI) have been examined by 1H nuclear magnetic resonance (n.m.r.). The intermediates were trapped during the reoxidation and consequent refolding of reduced PTI by alkylating free thiols; each intermediate contained different disulphide linkages. The n.m.r. spectra reveal that conformational features of the native protein are present in the intermediate containing just one of the three normal disulphide linkages (30-51). As additional normal disulphide bonds are formed, the conformation becomes more similar to that of the native protein. Introduction of additional but incorrect disulphide bonds does not lead to an increase in observable globular structure. A description of the folding process in terms of the conformations of the different intermediates is proposed. The significance of these results for the general mechanism of protein folding is outlined.

Amino Acid Sequence

A new two-disulphide intermediate in the refolding of reduced bovine pancreatic trypsin inhibitor.

The apparently complete refolding of reduced bovine pancreatic trypsin inhibitor (BPTI) is shown to produce a mixture of two species. One of these is native BPTI, but the other lacks the disulphide bond between cysteines 30 and 51. The latter species has a folded conformation very like that of native BPTI, and is oxidized by air to native BPTI on warming in aqueous solution. The two unreactive cysteine thiol groups appear to be buried in the interior of the molecule, which restricts access by reagents that can alkylate them or oxidize them to form the disulphide bond. The implications of this intermediate and its conformation for the understanding of protein folding are discussed.

Amino Acid Sequence

Dynamic filtering by two-dimensional 1H NMR with application to phage lambda repressor.

Flexible regions of proteins play an important role in catalysis, ligand binding, and macromolecular interactions. Because of its enhanced sensitivity to motional narrowing, two-dimensional coupling constant J-correlated 1H NMR may be used to observe these regions selectively. Dynamic filtering is an intrinsic feature of this experiment because cross-peak amplitude decays rapidly as linewidths approach the coupling constant. We demonstrate here the flexibility of the NH2-terminal arm of phage lambda repressor, which is thought to wrap around the double helix in the repressor-operator complex. The assignment of arm resonances is made possible by the construction of mutant repressor genes containing successive NH2-terminal deletions.

Bacteriophage lambda