PubMed Health⌕ Search

Biomedical subjects

S Pietrokovski

Publications and source records attributed to S Pietrokovski.

At least 19 recordsLinked to original sources

Consistency analysis of similarity between multiple alignments: prediction of protein function and fold structure from analysis of local sequence motifs.

A new method to analyze the similarity between multiply aligned protein motifs (blocks) was developed. It identifies sets of consistently aligned blocks. These are found to be protein regions of similar function and structure that appear in different contexts. For example, the Rossmann fold ligand-binding region is found similar to TIM barrel and methylase regions, various protein families are predicted to have a TIM-barrel fold and the structural relation between the ClpP protease and crotonase folds is identified from their sequence. Besides identifying local structure features, sequence similarity across short sequence-regions (less than 20 amino acid regions) also predicts structure similarity of whole domains (folds) a few hundred amino acid residues long. Most of these relations could not be identified by other advanced sequence-to-sequence or sequence-to-multiple alignments comparisons. We describe the method (termed CYRCA), present examples of our findings, and discuss their implications.

Adenosine Triphosphatases↗

Identification of new signaling components in the Drosophila genome sequence.

The availability of the complete sequence of the Drosophila genome and the assignment of putative reading frames, provides an opportunity to search for new members in families of proteins generating signaling cascades. The six major pathways that dictate patterning were examined: receptor tyrosine kinases, transforming growth factor beta (TGF beta), Wnt, Toll, Hedgehog and Notch. Several new components were identified for the first four pathways, including ligands, receptors, cytoplasmic components and transcription factors. Most notable is the identification of a vascular endothelial growth factor (VEGF) receptor tyrosine kinase, two insulin/insulin growth factor I (IGF I) receptors without cytoplasmic protein kinase domains, and a family of proteins similar to Rhomboid (a protein involved in cleavage of TGF alpha-like ligands). A new TGF beta family ligand, two new Wnts and a Frizzled receptor were also identified. Finally, for the Toll pathway, two new potential Spatzle-like ligands and two new receptors were identified.

Animals↗

Intein spread and extinction in evolution.

Inteins are selfish DNA elements found within coding regions. They are translated with their host protein, but then catalyze their own excision and the formation of a peptide bond between their flanking protein regions. Understanding what drives and selects inteins is relevant for assessing whether they have unidentified biological functions and whether they can invade and become established in new genes and organisms. Inteins are suggested to have been present and more common in the progenitors of eukaryotes and prokaryotes. In these cells, inteins had some beneficial function or had evolved from an unknown beneficial protein. Since then, this putative benefit has been lost and inteins are gradually becoming extinct. The proteins in which inteins are currently found are proposed to be proteins vital for the survival of the organism, where intein removal is most difficult.

Archaea↗

Doublecortin mutations cluster in evolutionarily conserved functional domains.

Mutations in the X-linked gene doublecortin ( DCX ) result in lissencephaly in males or subcortical laminar heterotopia ('double cortex') in females. Various types of mutation were identified and the sequence differences included nonsense, splice site and missense mutations throughout the gene. Recently, we and others have demonstrated that DCX interacts and stabilizes microtubules. Here, we performed a detailed sequence analysis of DCX and DCX-like proteins from various organisms and defined an evolutionarily conserved Doublecortin (DC) domain. The domain typically appears in the N-terminus of proteins and consists of two tandemly repeated 80 amino acid regions. In the large majority of patients, missense mutations in DCX fall within the conserved regions. We hypothesized that these repeats may be important for microtubule binding. We expressed DCX or DCLK (KIAA0369) repeats in vitro and in vivo. Our results suggest that the first repeat binds tubulin but not microtubules and enhances microtubule polymerization. To study the functional consequences of DCX mutations, we overexpressed seven of the reported mutations in COS7 cells and examined their effect on the microtubule cytoskeleton. The results demonstrate that some of the mutations disrupt microtubules. The most severe effect was observed with a tyrosine to histidine mutation at amino acid 125 (Y125H). Produced as a recombinant protein, this mutation disrupts microtubules in vitro at high molar concentration. The positions of the different mutations are discussed according to the evolutionarily defined DC-repeat motif. The results from this study emphasize the importance of DCX-microtubule interaction during normal and abnormal brain development.

Amino Acid Sequence↗

Increased coverage of protein families with the blocks database servers.

The Blocks Database WWW (http://blocks.fhcrc.org ) and Email (blocks@blocks.fhcrc.org ) servers provide tools to search DNA and protein queries against the Blocks+ Database of multiple alignments, which represent conserved protein regions. Blocks+ nearly doubles the number of protein families included in the database by adding families from the Pfam-A, ProDom and Domo databases to those from PROSITE and PRINTS. Other new features include improved Block Searcher statistics, searching with NCBI's IMPALA program and 3D display of blocks on PDB structures.

Amino Acid Sequence↗

Blocks-based methods for detecting protein homology.

The most highly conserved regions of proteins can be represented as blocks of aligned sequence segments, typically with multiple blocks for a given protein family. The Blocks Database World Wide Web (http://blocks.fhcrc.org) and e-mail (blocks@blocks. fhcrc.org) servers provide tools to search DNA and protein queries against the Blocks+ Database of multiple alignments. We describe features for detection of distant relationships using blocks. Blocks+ includes protein families from the PROSITE, Prints, Pfam-A, ProDom and Domo databases. Other features include searching Blocks+ with the BLIMPS and NCBI's IMPALA programs, sequence logos, phylogenetic trees, three-dimensional display of blocks on PDB structures, and a polymerase chain reaction (PCR) primer design strategy based on blocks.

Amino Acid Sequence↗

Isolation and characterization of a split B-type DNA polymerase from the archaeon Methanobacterium thermoautotrophicum deltaH.

We describe here the isolation and characterization of a B-type DNA polymerase (PolB) from the archaeon Methanobacterium thermoautotrophicum DeltaH. Uniquely, the catalytic domains of M. thermoautotrophicum PolB are encoded from two different genes, a feature that has not been observed as yet in other polymerases. The two genes were cloned, and the proteins were overexpressed in Escherichia coli and purified individually and as a complex. We demonstrate that both polypeptides are needed to form the active polymerase. Similar to other polymerases constituting the B-type family, PolB possesses both polymerase and 3'-5' exonuclease activities. We found that a homolog of replication protein A from M. thermoautotrophicum inhibits the PolB activity. The inhibition of DNA synthesis by replication protein A from M. thermoautotrophicum can be relieved by the addition of M. thermoautotrophicum homologs of replication factor C and proliferating cell nuclear antigen. The possible roles of PolB in M. thermoautotrophicum replication are discussed.

Amino Acid Sequence↗

Configuration of the catalytic GIY-YIG domain of intron endonuclease I-TevI: coincidence of computational and molecular findings.

I-TevI is a member of the GIY-YIG family of homing endonucleases. It is folded into two structural and functional domains, an N-terminal catalytic domain and a C-terminal DNA-binding domain, separated by a flexible linker. In this study we have used genetic analyses, computational sequence analysis andNMR spectroscopy to define the configuration of theN-terminal domain and its relationship to the flexible linker. The catalytic domain is an alpha/beta structure contained within the first 92 amino acids of the 245-amino acid protein followed by an unstructured linker. Remarkably, this structured domain corresponds precisely to the GIY-YIG module defined by sequence comparisons of 57 proteins including more than 30 newly reported members of the family. Although much of the unstructured linker is not essential for activity, residues 93-116 are required, raising the possibility that this region may adopt an alternate conformation upon DNA binding. Two invariant residues of the GIY-YIG module, Arg27 and Glu75, located in alpha-helices, have properties of catalytic residues. Furthermore, the GIY-YIG sequence elements for which the module is named form part of a three-stranded antiparallel beta-sheet that is important for I-TevI structure and function.

Amino Acid Sequence↗

New features of the Blocks Database servers.

Blocks are ungapped multiple sequence alignments representing conserved protein regions, and the Blocks Database consists of blocks from documented protein families. World Wide Web (http://www. blocks.fhcrc.org) and Email (blocks@blocks.fhcrc.org) servers provide tools for homology searching and for analyzing protein family relationships. New enhancements include a multiple alignment processor that extends the use of these tools to imported multiple alignments of families not present in the database and a PCR primer designer that implements a new strategy for gene isolation.

DNA Primers↗

Blocks+: a non-redundant database of protein alignment blocks derived from multiple compilations.

MOTIVATION: As databanks grow, sequence classification and prediction of function by searching protein family databases becomes increasingly valuable. The original Blocks Database, which contains ungapped multiple alignments for families documented in Prosite, can be searched to classify new sequences. However, Prosite is incomplete, and families from other databases are now available to expand coverage of the Blocks Database. RESULTS: To take advantage of protein family information present in several existing compilations, we have used five databases to construct Blocks+, a unified database that is built on the PROTOMAT/BLOSUM scoring model and that can be searched using a single algorithm for consistent sequence classification. The LAMA blocks-versus-blocks searching program identifies overlapping protein families, making possible a non-redundant hierarchical compilation. Blocks+ consists of all blocks derived from PROSITE, blocks from Prints not present in PROSITE, blocks from Pfam-A not present in PROSITE or Prints, and so on for ProDom and Domo, for a total of 1995 protein families represented by 8909 blocks, doubling the coverage of the original Blocks Database. A challenge for any procedure aimed at non-redundancy is to retain related but distinct families while discarding those that are duplicates. We illustrate how using multiple compilations can minimize this potential problem by examining the SNF2 family of ATPases, which is detectably similar to distinct families of helicases and ATPases. AVAILABILITY: http://blocks.fhcrc.org/

Adenosine Triphosphatases↗

Consensus-degenerate hybrid oligonucleotide primers for amplification of distantly related sequences.

We describe a new primer design strategy for PCR amplification of unknown targets that are related to multiply-aligned protein sequences. Each primer consists of a short 3' degenerate core region and a longer 5' consensus clamp region. Only 3-4 highly conserved amino acid residues are necessary for design of the core, which is stabilized by the clamp during annealing to template molecules. During later rounds of amplification, the non-degenerate clamp permits stable annealing to product molecules. We demonstrate the practical utility of this hybrid primer method by detection of diverse reverse transcriptase-like genes in a human genome, and by detection of C5DNA methyltransferase homologs in various plant DNAs. In each case, amplified products were sufficiently pure to be cloned without gel fractionation. This COnsensus-DEgenerate Hybrid Oligonucleotide Primer (CODEHOP) strategy has been implemented as a computer program that is accessible over the World Wide Web (http://blocks.fhcrc.org/codehop.html) and is directly linked from the BlockMaker multiple sequence alignment site for hybrid primer prediction beginning with a set of related protein sequences.

Amino Acid Sequence↗

Superior performance in protein homology detection with the Blocks Database servers.

The Blocks Database World Wide Web (http://www.blocks.fhcrc.org ) and Email (blocks@blocks.fhcrc.org) servers provide tools for the detection and analysis of protein homology based on alignment blocks representing conserved regions of proteins. During the past year, searching has been augmented by supplementation of the Blocks Database with blocks from the Prints Database, for a total of 4754 blocks from 1163 families. Blocks from both the Blocks and Prints Databases and blocks that are constructed from sequences submitted to Block Maker can be used for blocks-versus-blocks searching of these databases with LAMA, and for viewing logos and bootstrap trees. Sensitive searches of up-to-date protein sequence databanks are carried out via direct links to the MAST server using position-specific scoring matrices and to the BLAST and PSI-BLAST servers using consensus-embedded sequence queries. Utilizing the trypsin family to evaluate performance, we illustrate the superiority of blocks-based tools over expert pairwise searching or Hidden Markov Models.

Amino Acid Sequence↗

Modular organization of inteins and C-terminal autocatalytic domains.

Analysis of the conserved sequence features of inteins (protein "introns") reveals that they are composed of three distinct modular domains. The N-terminal (N) and C-terminal (C) domains are predicted to perform different parts of the autocatalytic protein splicing reaction. An optional endonuclease domain (EN) is shown to correspond to different types of homing endonucleases in different inteins. The N domain contains motifs predicted to catalyze the first steps of protein splicing, leading to the cleavage of the intein N terminus from its protein host. Intein N domain motifs are also found in C-terminal autocatalytic domains (CADs) present in hedgehog and other protein families. Specific residues in the N domain of intein and CADs are proposed to form a charge relay system involved in cleaving their N-termini. The intein C domain is apparently unique to inteins and contains motifs that catalyze the final protein splicing steps: ligation of the intein flanks and cleavage of its C terminus to release the free intein and spliced host protein. All intein EN domains known thus far have dodecapeptide (DOD, LAGLI-DADG) type homing endonuclease motifs. This work identifies an EN domain with an HNH homing-endonuclease motif and two new small inteins with no EN domains. One of these small inteins might be inactive or a "pseudo intein." The results suggest a modular architecture for inteins, clarify their origin and relationship to other protein families, and extend recent experimental findings on the functional roles of intein N, C, and EN motifs.

Amino Acid Sequence↗

Gene families: the taxonomy of protein paralogs and chimeras.

Ancient duplications and rearrangements of protein-coding segments have resulted in complex gene family relationships. Duplications can be tandem or dispersed and can involve entire coding regions or modules that correspond to folded protein domains. As a result, gene products may acquire new specificities, altered recognition properties, or modified functions. Extreme proliferation of some families within an organism, perhaps at the expense of other families, may correspond to functional innovations during evolution. The underlying processes are still at work, and the large fraction of human and other genomes consisting of transposable elements may be a manifestation of the evolutionary benefits of genomic flexibility.

Amino Acid Sequence↗