PubMed HealthSearch

Biomedical subjects

J G Henikoff

Publications and source records attributed to J G Henikoff.

8 recordsLinked to original sources

The Blocks database--a system for protein classification.

The Blocks Database contains multiple alignments of conserved regions in protein families. The database can be searched by e-mail and World Wide Web(WWW) servers (http://blocks.fhcrc.org/help) to classify protein and nucleotide sequences.

Amino Acid Sequence

Automated construction and graphical presentation of protein blocks from unaligned sequences.

Protein blocks consist of multiply aligned sequence segments that correspond to the most highly conserved regions of protein families. Typically, a set of related proteins has more than one region in common and their relationship can be represented as a series of ungapped blocks separated by unaligned regions. Blockmaker is an automated system available by electronic mail (blockmaker@howard.fhcrc.org) and the World Wide Web (http://www.blocks.fhcrc.org4) that finds blocks in a group of related protein sequences submitted by the user. It adapts and extends existing algorithms to make them useful to biologists looking for conserved regions in a group of related proteins sequences. Two sets of blocks are returned, one in which candidate blocks are detected using the MOTIF algorithm and the other using a Gibbs sampler algorithm that has been adapted for full automation. This use of two block-finding methods based on completely different principles provides a 'reality check,' whereby a block detected by both methods is considered to be correct. Resulting blocks can be displayed using the information-based 'sequence logo' method, adapted to incorporate sequence weights, which provides an intuitive visual description of both the residue and the conservation information at each position. Blocks generated by this system are useful in diverse applications, such as searching databases and designing degenerate PCR primers. As an example, blocks made from amino acid sequences related to Caenorhabditis elegans Tc1 transposase were used to search GenBank, revealing that several fish and amphibian genomic sequences harbor previously unreported Tc1 homologs.

Algorithms

Position-based sequence weights.

Sequence weighting methods have been used to reduce redundancy and emphasize diversity in multiple sequence alignment and searching applications. Each of these methods is based on a notion of distance between a sequence and an ancestral or generalized sequence. We describe a different approach, which bases weights on the diversity observed at each position in the alignment, rather than on a sequence distance measure. These position-based weights make minimal assumptions, are simple to compute, and perform well in comprehensive evaluations.

Amino Acid Sequence

Protein family classification based on searching a database of blocks.

The most highly conserved regions of proteins can be represented as "blocks" of locally aligned sequence segments. Previously, an automated system was introduced to generate a database of blocks that is searched for local similarities using a sequence query. Here, we describe a method for searching this database that can also reveal significant global similarities. Local and global alignments are scored independently, so they can be used in concert to infer homology. A set of 7082 diverse sequences not represented in the database provided queries for testing this approach. The resulting distributions of scores led to guidelines for interpretation of search data and to the classification of 289 uncatalogued sequences into known groups. Thirty-eight of these relationships appear to be new discoveries. We also show how searching a database of blocks can be used to detect repeated domains and to find distinct cross-family relationships that were missed in searches of sequence databases.

Animals

A computerized intervention to improve timing of outpatient follow-up: a multicenter randomized trial in patients treated with warfarin. National Consortium of Anticoagulation Clinics.

OBJECTIVE: To evaluate a computerized scheduling model that employs nonlinear optimization to recommend optimal follow-up intervals for patients taking warfarin. DESIGN: Randomized trial. SETTING: 5 anticoagulation clinics. PATIENTS/PARTICIPANTS: 620 patients expected to receive warfarin for > or = 6 weeks. INTERVENTIONS: Computer-generated recommendations for scheduling the next visit were presented to or withheld from practitioners. MEASUREMENTS AND MAIN RESULTS: The main outcome measures were the follow-up interval scheduled by the provider, the interval at which the patient actually returned to clinic, and the quality of anticoagulation control (computed as the absolute difference between the measured and target prothrombin times [PTRs] or international normalized ratios [INRs]). Follow-up intervals scheduled for the patients whose practitioners received computer-generated recommendations were significantly longer than those for control patients (mean, 4.4 vs 3.5 weeks, p < 0.001), despite the fact that the practitioners modified the suggested return interval by > 1 week on 40% of the visits. The interval at which the intervention group actually returned to clinic was also longer (mean, 4.4 vs 4.1 weeks, p < 0.05), even though the control patients tended to return at longer intervals than were scheduled by their practitioners. Control of anticoagulation was nearly the same among experimental and control patients. Life-threatening complications occurred in the care of three experimental patients and one control patient, while other serious complications occurred in the care of 16 experimental patients and 17 control patients. CONCLUSIONS: Recommendations based on nonlinear optimization prompted clinicians to schedule less frequent follow-up for patients taking warfarin, with no deterioration in anticoagulation control. This approach to scheduling can potentially reduce utilization while maintaining quality of care for patients who require long-term monitoring.

Appointments and Schedules

Performance evaluation of amino acid substitution matrices.

Several choices of amino acid substitution matrices are currently available for searching and alignment applications. These choices were evaluated using the BLAST searching program, which is extremely sensitive to differences among matrices, and the Prosite catalog, which lists members of hundreds of protein families. Matrices derived directly from either sequence-based or structure-based alignments of distantly related proteins performed much better overall than extrapolated matrices based on the Dayhoff evolutionary model. Similar results were obtained with the FASTA searching program. Improved performance appears to be general rather than family-specific, reflecting improved accuracy in scoring alignments. An implementation of a multiple matrix strategy was also tested. While no combination of three matrices performed as well as the single best matrix, BLOSUM 62, good results were obtained using a combination of sequence-based and structure-based matrices. This hybrid set of matrices is likely to be useful in certain situations. Our results illustrate the importance of matrix selection and the value of a comprehensive approach to evaluation of protein comparison tools.

Amino Acid Sequence

Amino acid substitution matrices from protein blocks.

Methods for alignment of protein sequences typically measure similarity by using a substitution matrix with scores for all possible exchanges of one amino acid with another. The most widely used matrices are based on the Dayhoff model of evolutionary rates. Using a different approach, we have derived substitution matrices from about 2000 blocks of aligned sequence segments characterizing more than 500 groups of related proteins. This led to marked improvements in alignments and in searches using queries from each of the groups.

Algorithms

Automated assembly of protein blocks for database searching.

A system is described for finding and assembling the most highly conserved regions of related proteins for database searching. First, an automated version of Smith's algorithm for finding motifs is used for sensitive detection of multiple local alignments. Next, the local alignments are converted to blocks and the best set of non-overlapping blocks is determined. When the automated system was applied successively to all 437 groups of related proteins in the PROSITE catalog, 1764 blocks resulted; these could be used for very sensitive searches of sequence databases. Each block was calibrated by searching the SWISS-PROT database to obtain a measure of the chance distribution of matches, and the calibrated blocks were concatenated into a database that could itself be searched. Examples are provided in which distant relationships are detected either using a set of blocks to search a sequence database or using sequences to search the database of blocks. The practical use of the blocks database is demonstrated by detecting previously unknown relationships between oxidoreductases and by evaluating a proposed relationship between HIV Vif protein and thiol proteases.

Algorithms