PubMed HealthSearch

Biomedical subjects

H Margalit

Publications and source records attributed to H Margalit.

At least 19 recordsLinked to original sources

Molecular characterization of a common fragile site (FRA7H) on human chromosome 7 by the cloning of a simian virus 40 integration site.

Common fragile sites are chromosomal loci prone to breakage and rearrangement, hypothesized to provide targets for foreign DNA integration. We cloned a simian virus 40 integration site and showed by fluorescent in situ hybridization analysis that the integration event had occurred within a common aphidicolin-induced fragile site on human chromosome 7, FRA7H. A region of 161 kb spanning FRA7H was defined and sequenced. Several regions with a potential unusual DNA structure, including high-flexibility, low-stability, and non-B-DNA-forming sequences were identified in this region. We performed a similar analysis on the published FRA3B sequence and the putative partial FRA7G, which also revealed an impressive cluster of regions with high flexibility and low stability. Thus, these unusual DNA characteristics are possibly intrinsic properties of common fragile sites that may affect their replication and condensation as well as organization, and may lead to fragility.

Base Sequence

Quantitative parameters for amino acid-base interaction: implications for prediction of protein-DNA binding sites.

Inspection of the amino acid-base interactions in protein-DNA complexes is essential to the understanding of specific recognition of DNA target sites by regulatory proteins. The accumulation of information on protein-DNA co-crystals challenges the derivation of quantitative parameters for amino acid-base interaction based on these data. Here we use the coordinates of 53 solved protein-DNA complexes to extract all non-homologous pairs of amino acid-base that are in close contact, including hydrogen bonds and hydrophobic interactions. By comparing the frequency distribution of the different pairs to a theoretical distribution and calculating the log odds, a quantitative measure that expresses the likelihood of interaction for each pair of amino acid-base could be extracted. A score that reflects the compatibility between a protein and its DNA target can be calculated by summing up the individual measures of the pairs of amino acid-base involved in the complex, assuming additivity in their contributions to binding. This score enables ranking of different DNA binding sites given a protein binding site and vice versa and can be used in molecular design protocols. We demonstrate its validity by comparing the predictions using this score with experimental binding results of sequence variants of zif268 zinc fingers and their DNA binding sites.

Amino Acids

A role for CH...O interactions in protein-DNA recognition.

The concept of CH...O hydrogen bonds has recently gained much interest, with a number of reports indicating the significance of these non-classical hydrogen bonds in stabilizing nucleic acid and protein structures. Here, we analyze the CH...O interactions in the protein-DNA interface, based on 43 crystal structures of protein-DNA complexes. Surprisingly, we find that the number of close intermolecular CH...O contacts involving the thymine methyl group and position C5 of cytosine is comparable to the number of protein-DNA hydrogen bonds involving nitrogen and oxygen atoms as donors and acceptors. A comprehensive analysis of the geometries of these close contacts shows that they are similar to other CH...O interactions found in proteins and small molecules, as well as to classical NH...O hydrogen bonds. Thus, we suggest that C5 of cytosine and C5-Met of thymine form relatively weak CH...O hydrogen bonds with Asp, Asn, Glu, Gln, Ser, and Thr, contributing to the specificity of recognition. Including these interactions, in addition to the classical protein-DNA hydrogen bonds, enables the extraction of simple structural principles for amino acid-base recognition consistent with electrostatic considerations.

Base Composition

A structure-based algorithm to predict potential binding peptides to MHC molecules with hydrophobic binding pockets.

Binding of peptides to MHC class I molecules is a prerequisite for their recognition by cytotoxic T cells. Consequently, identification of peptides that will bind to a given MHC molecule must constitute a central part of any algorithm for prediction of T-cell antigenic peptides based on the amino acid sequence of the protein. Binding motifs, defined by anchor positions only, have proven to be insufficient to ensure binding, suggesting that other positions along the peptide sequence also affect peptide-MHC interaction. The second phase of prediction schemes therefore take into account the effect of all positions along the peptide sequence, and are based on position-dependent-coefficients that are used in the calculation of a peptide score. These coefficients can be extracted from a large ensemble of binding sequences that were tested experimentally, or derived from structural considerations, as in the algorithm developed by us recently. This algorithm uses the coordinates of solved complexes to evaluate the interactions of peptide amino acids with MHC contact residues, and results in a peptide score that reflects its binding energy. Here we present our analysis for peptide binding to four MHC alleles (HLA-A2, HLA-A68, HLA-B27 and H-2Kb), and compare the predictions of the algorithm to experimental binding data. The algorithm performs successfully in predicting peptide binding to MHC molecules with hydrophobic binding pockets but not when MHC molecules with hydrophilic, charged pockets are considered. For MHC molecules with hydrophobic pockets it is demonstrated how the algorithm succeeds in distinguishing binding from non-binding peptides, and in high ranking of immunogenic peptides within all overlapping same-length peptides spanning their respective protein sequences. The latter property of the algorithm makes it a useful tool in the rational design of peptide vaccines aimed at T-cell immunity.

Algorithms

Microsatellite spreading in the human genome: evolutionary mechanisms and structural implications.

Microsatellites are tandem repeat sequences abundant in the genomes of higher eukaryotes and hitherto considered as "junk DNA." Analysis of a human genome representative data base (2.84 Mb) reveals a distinct juxtaposition of A-rich microsatellites and retroposons and suggests their coevolution. The analysis implies that most microsatellites were generated by a 3'-extension of retrotranscripts, similar to mRNA polyadenylylation, and that they serve in turn as "retroposition navigators," directing the retroposons via homology-driven integration into defined sites. Thus, they became instrumental in the preservation and extension of primordial genomic patterns. A role is assigned to these reiterating A-rich loci in the higher-order organization of the chromatin. The disease-associated triplet repeats are mostly found in coding regions and do not show an association with retroposons, constituting a unique set within the family of microsatellite sequences.

Biological Evolution

Structure and function of the Pseudomonas putida integration host factor.

Integration host factor (IHF) is a DNA-binding and -bending protein that has been found in a number of gram-negative bacteria. Here we describe the cloning, sequencing, and functional analysis of the genes coding for the two subunits of IHF from Pseudomonas putida. Both the ihfA and ihfB genes of P. putida code for 100-amino-acid-residue polypeptides that are 1 and 6 residues longer than the Escherichia coli IHF subunits, respectively. The P. putida ihfA and ihfB genes can effectively complement E. coli ihf mutants, suggesting that the P. putida IHF subunits can form functional heterodimers with the IHF subunits of E. coli. Analysis of the amino acid differences between the E. coli and P. putida protein sequences suggests that in the evolution of IHF, amino acid changes were mainly restricted to the N-terminal domains and to the extreme C termini. These changes do not interfere with dimer formation or with DNA recognition. We constructed a P. putida mutant strain carrying an ihfA gene knockout and demonstrated that IHF is essential for the expression of the P(U) promoter of the xyl operon of the upper pathway of toluene degradation. It was further shown that the ihfA P. putida mutant strain carrying the TOL plasmid was defective in the degradation of the aromatic model compound benzyl alcohol, proving the unique role of IHF in xyl operon promoter regulation.

Amino Acid Sequence

Comprehensive analysis of hydrogen bonds in regulatory protein DNA-complexes: in search of common principles.

A systematic analysis of hydrogen bonds between regulatory proteins and their DNA targets is presented, based on 28 crystallographically solved complexes. All possible hydrogen bonds were screened and classified into different types: those that involve the amino acid side-chains and DNA base edges and those that involve the backbone atoms of the molecules. For each interaction type, all bonds were characterized and a statistical analysis was performed to reveal significant amino acid-base interdependence. The interactions between the amino acid side-chains and DNA backbone constitute about half of the interactions, but did not show any amino acid-base correlation. Interactions via the protein backbone were also observed, predominantly with the DNA backbone. As expected, the most significant pairing preference was demonstrated for interactions between the amino acid side-chains and the DNA base edges. The statistically significant relationships could mostly be explained by the chemical nature of the participants. However, correlations that could not be trivially predicted from the hydrogen bonding potential of the residues were also identified, like the preference of lysine for guanine over adenine, or the preference of glutamic acid for cystosine over adenine. While Lys x G interactions were very frequent and spread over various families, the Glu x C interactions were found mainly in the basic helix-loop-helix family. Further examination of the side-chain-base edge contacts at the atomic level revealed a trend of the amino acids to contact the DNA by their donor atoms, preferably at position W2 in the major groove. In most cases it seems that the interactions are not guided simply by the presence of a required atom in a specific position in the groove, but that the identity of the base possessing this atom is crucial. This may have important implications in molecular design experiments.

Amino Acids

Periodic variation in side-chain polarities of T-cell antigenic peptides correlates with their structure and activity.

We present an analysis that synthesizes information on the sequence, structure, and motifs of antigenic peptides, which previously appeared to be in conflict. Fourier analysis of T-cell antigenic peptides indicates a periodic variation in amino acid polarities of 3-3.6 residues per period, suggesting an amphipathic alpha-helical structure. However, the diffraction patterns of major histocompatibility complex (MHC) molecules indicate that their ligands are in an extended non-alpha-helical conformation. We present two mutually consistent structural explanations for the source of the alpha-helical periodicity, based on an observation that the side chains of MHC-bound peptides generally partition with hydrophobic (hydrophilic) side chains pointing into (out of) the cleft. First, an analysis of haplotype-dependent peptide motifs indicates that the locations of their defining residues tend to force a period 3-4 variation in hydrophobicity along the peptide sequence, in a manner consistent with the spacing of pockets in the MHC. Second, recent crystallographic determination of the structure of a peptide bound to a class II MHC molecule reveals an extended but regularly twisted peptide with a rotation angle of about 130 degrees. We show that similar structures with rotation angles of 100-130 degrees are energetically acceptable and also span the length of the MHC cleft. These results provide a sound physical chemical and structural basis for the existence of a haplotype-independent antigenic motif which can be particularly important in limiting the search time for antigenic peptides.

Antigens

Ranking potential binding peptides to MHC molecules by a computational threading approach.

In this paper, an approach developed to address the inverse protein folding problem is applied to prediction of potential binding peptides to a specific major histocompatibility complex (MHC) molecule. Overlapping peptides, spanning the entire protein sequence, are threaded through the backbone coordinates of a known peptide fold in the MHC groove, and their interaction energies are evaluated using statistical pairwise contact potentials. With currently available tables for pairwise potentials, promising results are obtained for MHC-peptide complexes where hydrophobic interactions predominate. By ranking the peptides in an ascending order according to their energy values, it is demonstrated that, in most cases, known antigenic peptides are highly ranked. Furthermore, predicted hierarchies are consistent with experimental binding results. Currently, predictions of potential binding peptides to a specific MHC molecule are based on the identification of allele-specific binding motifs. However, it has been demonstrated that these motifs are neither sufficient nor strictly required to ensure binding. The computational procedure presented here succeeds in determining the MHC binding potential of peptides along a protein amino acid sequence, without relying on binding motifs. The proposed scheme may significantly reduce the number of peptides to be tested, identify good binders that do not necessarily show the known allele-specific binding motifs, and identify the best candidates among those with the motifs. In general, when structural information about a protein-peptide complex is available, the current application of the threading approach can be used to screen a large library of peptides for selection of the best binders to the target protein.

Amino Acid Sequence

Conservation of salt bridges in protein families.

A detailed computational analysis is presented that focuses on the relationship between structural attributes and the degree and mode of salt bridge conservation. A data set of conserved and non-conserved salt bridges was constructed from eight protein families, based on the structural alignment of family members. Salt bridges were defined at the secondary structure level rather than at the residue level, implying different possible modes of conservation: preservation (same charges at the same residue positions), compensation (reversal of charges), and complementation (maintenance of a salt bridge between two segments of secondary structures, not involving the same residue positions). Structural attributes such as the surface accessibility, distance from the active site, or type of secondary structures involved, were studied. No significant differences were found between conserved and non-conserved salt bridges, except for the surface accessibility. Conserved salt bridges were shown to be less exposed than non-conserved ones. Moreover, within the set of conserved salt bridges, the degree of conservation was shown to negatively correlate with surface exposure; however, not to an extent that could indicate a general role for electrostatic interactions in the protein interior. Examination of the most conserved salt bridge in each family showed a variety of typical features: Some involved the terminal segments of the protein, some were buried and one involved the catalytic site of the protein. Hence, the role of salt bridges is more specific, probably in fine tuning of a specific structure through the folding process or in determining the functional site. As for the conservation mode, preservations were found to predominate in the conserved interactions, while complementations were of secondary importance. Compensations occurred only rarely and mostly in exposed salt bridges, suggesting that this mechanism is not utilized frequently and especially not in important interactions.

Binding Sites

Identification of common motifs in unaligned DNA sequences: application to Escherichia coli Lrp regulon.

We describe a relatively simple method for the identification of common motifs in DNA sequences that are known to share a common function. The input sequences are unaligned and there is no information regarding the position or orientation of the motif. Often such data exists for protein-binding regions, where genetic or molecular information that defines the binding region is available, but the specific recognition site within it is unknown. The method is based on the principle of 'divide and conquer'; we first search for dominant submotifs and then build full-length motifs around them. This method has several useful features: (i) it screens all submotifs so that the results are independent of the sequence order in the data; (ii) it allows the submotifs to contain spacers; (iii) it identifies an existing motif even if the data contains 'noise'; (iv) its running time depends linearly on the total length of the input. The method is demonstrated on two groups of protein-binding sequences: a well-studied group of known CRP-binding sequences, and a relatively newly identified group of genes known to be regulated by Lrp. The Lrp motif that we identify, based on 23 gene sequences, is similar to a previously identified motif based on a smaller data set, and to a consensus sequence of experimentally defined binding sites. Individual Lrp sites are evaluated and compared in regard to their regulation mode.

Algorithms

Determination of common structural features in Escherichia coli promoters by computer analysis.

Escherichia coli promoters show a large degree of sequence variation. However, they are all recognized specifically by RNA polymerase as the sites for transcription initiation, suggesting that they share common basic structural features distinguishing them from the rest of the sequence. Our hypothesis is that the promoter is determined not only by the two consensus sequences at -10 and -35, but also by the surrounding nucleotides, and that it is not only the identity of the nucleotides that is important for promoter function but the presence of specific physical-chemical and structural characteristics that are sequence dependent. This approach is supported by accumulating evidence indicating the role that the DNA conformation may play in modulating protein-DNA interaction. In this study, four intrinsic sequence-dependent characteristics are examined in E. coli promoter regions: helix stability, helix flexibility, and two conformational parameters represented by the DNA tendencies for B-->Z and B-->A transition. The promoter is defined by the consensus sequences and their vicinity and the examined properties are compared between promoter and random sequences. It is demonstrated that both the consensus and flanking regions are less stable, more flexible and show a higher tendency for the B conformation in comparison to random sequences. Discriminant analysis is used to evaluate the relative contributions of the various characteristics.

Databases, Factual

Sequence features that correlate with MHC restriction.

Identification of common sequence motifs in antigenic peptides restricted to a specific class II molecule has not been easy due to the large variation in length and sequence that is observed in these peptides. The goal of this study is to develop an automated computerized method for the identification of sequence features and structural determinants that play a role in the MHC restriction of helper T-cell antigenic peptides. For this, we compiled an extended database of helper T-cell sites, including the information on MHC restriction, when available. Two groups of peptides are assigned to each MHC type: (1) peptides that bind to that MHC molecule to elicit a T-cell response, and (2) peptides that were shown experimentally either not to bind to or not to elicit a T-cell proliferative response in association with that MHC molecule. We search for common motifs in the group of binding peptides, and identify significant motifs that are frequent among these peptides but almost absent in the group of non-binding peptides. A motif consists of physical-chemical and structural properties that may be responsible for binding specificity and can be extracted from sequence data, such as, hydrophobicity, charge, hydrogen bonding capability, etc. The first search is performed on the non-aligned binding peptides. Next, the sequences are aligned according to an identified motif and a search for additional, conserved, properties is performed. The statistical significance of the motifs is evaluated as well as their compatibility with published experimental results on substitution effects. Here we demonstrate the general scheme of the analysis and results for I-Ek and I-Ak associated peptides.

Amino Acid Sequence

Comparative analysis of structurally defined heparin binding sequences reveals a distinct spatial distribution of basic residues.

Heparin, among the best studied glycosaminoglycans, is well known for its involvement in a variety of physiological processes. Many proteins, whose activities are modulated via heparin binding, were identified, and the consequences of their interaction with heparin were characterized. However, in the absence of solid structural information regarding heparin-protein complexes, the mechanism by which heparin operates at the molecular level is still obscure. The structure of such a complex is hereby explored via the identification of a common motif in heparin binding sequences. To avoid ambiguity we included in our data base only sequences that have been shown experimentally to be directly involved in heparin binding. Then, a comparison of the spatial distribution of basic residues was conducted among those peptides for which three-dimensional structures were defined. Using computer graphics techniques we were able to identify a unique distribution shared by all of these segments. Two basic amino acids (most frequently arginine) are located at about 20 A apart, facing opposite directions of an alpha-helix. Other basic amino acids are dispersed between these two residues, facing one side, while nonpolar residues face the opposite side, forming an amphipathic structure. The distribution of basic amino acids in other heparin binding sequences that preserves the same spatial arrangement seems to be compatible with a beta-strand structure. The 20-A interval accommodates a glycosaminoglycan pentasaccharide, and the spatial distribution of the basic residues suggests an intertwinement of the heparin-protein complex. The dynamics of such an interaction may provide a clue regarding the ensuing change in protein activity.

Amino Acid Sequence

Identification and characterization of E.coli ribosomal binding sites by free energy computation.

Sequences upstream from translational initiation sites of different E.coli genes show various degrees of complementarity to the Shine-Dalgarno (SD) sequence at the 3' end of the 16S rRNA. We propose a quantitative measure for the SD region on the mRNA, that reflects its degree of complementarity to the rRNA. This measure is based on the stability of the rRNA-mRNA duplex as established by free energy computations. The free energy calculations are based on the same principles that are used for folding a single RNA molecule, and are executed by similar algorithms. Bulges and internal loops in the rRNA and mRNA are allowed. The mRNA string with maximum free energy gain upon binding to the rRNA is selected as the most favorable SD sequence of a gene. The free energy value that represents the SD region provides a quantitative measure that can be used for comparing SD sequences of different genes. The distribution of this measure in more than 1000 E.coli genes is presented and discussed.

Base Sequence

Compilation of E. coli mRNA promoter sequences.

An updated compilation of 300 E. coli mRNA promoter sequences is presented. For each sequence the most recent relevant paper was checked, to verify the location of the transcriptional start position as identified experimentally. We comment on the reliability of the sequence databanks and analyze the conservation of known promoter features in the current compilation. This database is available by E-mail.

Base Sequence

Molecular analysis of HLA class II genes in primary Sjögren's syndrome. A study of Israeli Jewish and Greek non-Jewish patients.

In an attempt to define the role of HLA class II genes in predisposition to primary Sjögren's syndrome, patients of two different ethnic groups (Israeli Jews and Greeks of non-Jewish origin) suffering from this disorder were studied. Oligonucleotide genotyping revealed the majority in both groups to carry either DRB1*1101 or DRB1*1104, alleles that are in linkage disequilibrium with DQB1*0301 and DQA1*0501. The high frequency of the two alleles in these SS patients is in contrast with the accepted association of primary SS with HLA-DR3 in Italian and American individuals. Molecular analysis of DQB1 and DQA1 alleles found in American Caucasian and American black SS (or SLE) patients demonstrated high frequencies of DQB1*0201 and DQA1*0501. The fact that the majority of SS patients, across racial and ethnic boundaries, carry a common allele, DQA1*0501, implies its involvement in the predisposition to primary SS. Based on sequence analysis and the computer imaging of the HLA class II molecule structure, a hypothetical model for the role of the DQ molecule in promoting primary SS is proposed.

Alleles