PubMed HealthSearch

Biomedical subjects

S Henikoff

Publications and source records attributed to S Henikoff.

At least 19 recordsLinked to original sources

Automated construction and graphical presentation of protein blocks from unaligned sequences.

Protein blocks consist of multiply aligned sequence segments that correspond to the most highly conserved regions of protein families. Typically, a set of related proteins has more than one region in common and their relationship can be represented as a series of ungapped blocks separated by unaligned regions. Blockmaker is an automated system available by electronic mail (blockmaker@howard.fhcrc.org) and the World Wide Web (http://www.blocks.fhcrc.org4) that finds blocks in a group of related protein sequences submitted by the user. It adapts and extends existing algorithms to make them useful to biologists looking for conserved regions in a group of related proteins sequences. Two sets of blocks are returned, one in which candidate blocks are detected using the MOTIF algorithm and the other using a Gibbs sampler algorithm that has been adapted for full automation. This use of two block-finding methods based on completely different principles provides a 'reality check,' whereby a block detected by both methods is considered to be correct. Resulting blocks can be displayed using the information-based 'sequence logo' method, adapted to incorporate sequence weights, which provides an intuitive visual description of both the residue and the conservation information at each position. Blocks generated by this system are useful in diverse applications, such as searching databases and designing degenerate PCR primers. As an example, blocks made from amino acid sequences related to Caenorhabditis elegans Tc1 transposase were used to search GenBank, revealing that several fish and amphibian genomic sequences harbor previously unreported Tc1 homologs.

Algorithms

Conservation of brown gene trans-inactivation in Drosophila.

The mechanism underlying trans-inactivation associated with dominant position effect variegation (PEV) of the Drosophila melanogaster brown gene has been addressed by a comparison with its D. virilis homologue. This comparison revealed: 86% identity between conceptual translation products of the brown gene from these two species, functional homology, as the D. virilis gene rescues a D. melanogaster null brown mutation, and conservation of the sequences required for trans-inactivation, as the D. virilis gene in D. melanogaster is subject to dominant PEV. An extended region of sequence similarity upstream of the open reading frame is observed. As the D. virilis homologue is functionally interchangeable with the D. melanogaster gene, these genes must share regulatory sequences as well as protein coding homology. These results support a model in which trans-inactivation is mediated by a heterochromatin-sensitive transcription factor.

ATP-Binding Cassette Transporters

Distance and pairing effects on the brownDominant heterochromatic element in Drosophila.

We examined the behavior of the brownDominant (bwD) heterochromatic insertion moved to different locations relative to centric heterochromatin. Effects were measured as the degree of silencing of a wild-type brown eye pigment gene by bwD across a tandem duplication. A series of X-ray-induced effects were recovered at high frequency. Cis-acting enhancers were obtained by relocation of the duplication closer to autosomal heterochromatin. Enhancers were also recovered on the homologous chromosome when it was similarly rearranged, revealing a novel interhomologue effect whereby interactions occur between genetic elements near opposite ends of a chromosome arm rather than between paired alleles. Cis-acting suppressors were obtained as secondary rearrangements in which the duplication was moved farther away from heterochromatin. Suppression was correlated with loss of cytological association between bwD and the polytene chromocenter. Surprisingly, the distance from bwD to the chromocenter was not correlated with the strength of enhancement or suppression. We propose that bwD fails to coalesce with the chromocenter when its position along the chromosome places it beyond a threshold distance from heterochromatin, and this threshold depends upon the configuration of both the chromosome carrying bwD and its paired homologue.

Alleles

Position-based sequence weights.

Sequence weighting methods have been used to reduce redundancy and emphasize diversity in multiple sequence alignment and searching applications. Each of these methods is based on a notion of distance between a sequence and an ancestral or generalized sequence. We describe a different approach, which bases weights on the diversity observed at each position in the alignment, rather than on a sequence distance measure. These position-based weights make minimal assumptions, are simple to compute, and perform well in comprehensive evaluations.

Amino Acid Sequence

Expansions of transgene repeats cause heterochromatin formation and gene silencing in Drosophila.

Closely linked repeats of a Drosophila P transposon carrying a white transgene were found to cause white variegation. Arrays of three or more transgenes produced phenotypes similar to classical heterochromatin-induced position-effect variegation (PEV), and these phenotypes were modified by known modifiers of PEV. This effect on the repeated transgenes was much stronger for a site near centric heterochromatin than it was for a medial site, and it strengthened with increasing copy number. Differences between variegated phenotypes could be accounted for if different topological structures were generated by pairing between closely linked repeat sequences. We propose that pairing of repeats underlies heterochromatin formation and is responsible for diverse gene silencing phenomena in animals and plants.

Animals

Protein family classification based on searching a database of blocks.

The most highly conserved regions of proteins can be represented as "blocks" of locally aligned sequence segments. Previously, an automated system was introduced to generate a database of blocks that is searched for local similarities using a sequence query. Here, we describe a method for searching this database that can also reveal significant global similarities. Local and global alignments are scored independently, so they can be used in concert to infer homology. A set of 7082 diverse sequences not represented in the database provided queries for testing this approach. The resulting distributions of scores led to guidelines for interpretation of search data and to the classification of 289 uncatalogued sequences into known groups. Thirty-eight of these relationships appear to be new discoveries. We also show how searching a database of blocks can be used to detect repeated domains and to find distinct cross-family relationships that were missed in searches of sequence databases.

Animals

Modification of the Drosophila heterochromatic mutation brownDominant by linkage alterations.

The variegating mutation brownDominant (bwD) of Drosophila melanogaster is associated with an insertion of heterochromatin into chromosome arm 2R at 59E, the site of the bw gene. Mutagenesis produced 150 dominant suppressors of bwD variegation. These fall into two classes: unlinked suppressors, which also suppress other variegating mutations; and linked chromosome rearrangements, which suppress only bwD. Some rearrangements are broken at 59E, and so might directly interfere with variegation caused by the heterochromatic insertion at that site. However, most rearrangements are translocations broken proximal to bw within the 52D-57D region of 2R. Translocation breakpoints on the X chromosome are scattered throughout the X euchromatin, while those on chromosome 3 are confined to the tips. This suggests that a special property of the X chromosome suppresses bwD variegation, as does a distal autosomal location. Conversely, two enhancers of bwD are caused by translocations from the same part of 2R to proximal heterochromatin, bringing the bwD heterochromatic insertion close to the chromocenter with which it strongly associates. These results support the notion that heterochromatin formation at a genetic locus depends on its location within the nucleus.

Animals

Performance evaluation of amino acid substitution matrices.

Several choices of amino acid substitution matrices are currently available for searching and alignment applications. These choices were evaluated using the BLAST searching program, which is extremely sensitive to differences among matrices, and the Prosite catalog, which lists members of hundreds of protein families. Matrices derived directly from either sequence-based or structure-based alignments of distantly related proteins performed much better overall than extrapolated matrices based on the Dayhoff evolutionary model. Similar results were obtained with the FASTA searching program. Improved performance appears to be general rather than family-specific, reflecting improved accuracy in scoring alignments. An implementation of a multiple matrix strategy was also tested. While no combination of three matrices performed as well as the single best matrix, BLOSUM 62, good results were obtained using a combination of sequence-based and structure-based matrices. This hybrid set of matrices is likely to be useful in certain situations. Our results illustrate the importance of matrix selection and the value of a comprehensive approach to evaluation of protein comparison tools.

Amino Acid Sequence

Amino acid substitution matrices from protein blocks.

Methods for alignment of protein sequences typically measure similarity by using a substitution matrix with scores for all possible exchanges of one amino acid with another. The most widely used matrices are based on the Dayhoff model of evolutionary rates. Using a different approach, we have derived substitution matrices from about 2000 blocks of aligned sequence segments characterizing more than 500 groups of related proteins. This led to marked improvements in alignments and in searches using queries from each of the groups.

Algorithms

A relationship between asparagine synthetase A and aspartyl tRNA synthetase.

A highly conserved protein motif characteristic of Class II aminoacyl tRNA synthetases was found to align with a region of Escherichia coli asparagine synthetase A. The alignment was most striking for aspartyl tRNA synthetase, an enzyme with catalytic similarities to asparagine synthetase. To test whether this sequence reflects a conserved function, site-directed mutagenesis was used to replace the codon for Arg298 of asparagine synthetase A, which aligns with an invariant arginine in the Class II aminoacyl tRNA synthetases. The resulting genes were expressed in E. coli, and the gene products were assayed for asparagine synthetase activity in vitro. Every substitution of Arg298, even to a lysine, resulted in a loss of asparagine synthetase activity. Directed random mutagenesis was then used to create a variety of codon changes which resulted in amino acid substitutions within the conserved motif surrounding Arg298. Of the 15 mutant enzymes with amino acid substitutions yielding soluble enzyme, 13 with changes within the conserved region were found to have lost activity. These results are consistent with the possibility that asparagine synthetase A, one of the two unrelated asparagine synthetases in E. coli, evolved from an ancestral aminoacyl tRNA synthetase.

Amino Acid Sequence