PubMed Health⌕ Search

Biomedical subjects

David Landsman

Publications and source records attributed to David Landsman.

6 recordsLinked to original sources

Characterization of sequence variability in nucleosome core histone folds.

The three-helix, approximately 65-residue histone fold domain is the most structurally conserved part of the core histones H2A, H2B, H3, and H4. However, it evinces a notable degree of sequence variation within and between histone classes. We used two approaches to characterize sequence variation in these histone folds, toward elucidating their structure/function relationships and evolution. On the one hand we asked how much of the sequence variation seen in structure-based alignments of the folds maintains physicochemical properties at a position, and on the other, whether conservation correlates to structural importance, as measured by the number of residue-to-residue contacts a position makes. Strong physicochemical conservation or correlation of conservation to contacts would support the idea that functional constraints, rather than genetic drift, determines the observed range of variants at a given position. We used an 11-state table of physicochemical properties to classify each position in the core histone fold (CHF) alignments, and a public website (http://www.ebi.ac.uk/thornton-srv/databases/cgi-bin/valdar/scorecons_server.pl) to score conservation. We found that, depending on histone class, from 38 to 77% of CHF positions are maximally conserved physicochemically, and that for H2B, H3, and H4 the degree to which a position is conserved correlates positively to the number of contacts made by the residue at that position in the crystal structure of the nucleosome core particle. We also examined the correlation between conservation and the type of contact (e.g., inter- or intrachain, histone-histone, or histone-DNA, etc.). For H2B, H3, and H4 we found a positive correlation between conservation and number of interchain protein contacts. No such correlation or statistical significance was found for DNA or intrachain contacts. This suggests that variations in the CHF sequences could be functionally constrained by requirements to make sufficient interchain histone contacts. We also suggest that inventory of histone residue variants can augment functional studies of histones. An example is presented for histone H3.

Binding Sites↗

Identifying related L1 retrotransposons by analyzing 3' transduced sequences.

BACKGROUND: A large fraction of the human genome is attributable to L1 retrotransposon sequences. Not only do L1s themselves make up a significant portion of the genome, but L1-encoded proteins are thought to be responsible for the transposition of other repetitive elements and processed pseudogenes. In addition, L1s can mobilize non-L1, 3'-flanking DNA in a process called 3' transduction. Using computational methods, we collected DNA sequences from the human genome for which we have high confidence of their mobilization through L1-mediated 3' transduction. RESULTS: The precursors of L1s with transduced sequence can often be identified, allowing us to reconstruct L1 element families in which a single parent L1 element begot many progeny L1s. Of the L1s exhibiting a sequence structure consistent with 3' transduction (L1 with transduction-derived sequence, L1-TD), the vast majority were located in duplicated regions of the genome and thus did not necessarily represent unique insertion events. Of the remaining L1-TDs, some lack a clear polyadenylation signal, but the alignment between the parent-progeny sequences nevertheless ends in an A-rich tract of DNA. CONCLUSIONS: Sequence data suggest that during the integration into the genome of RNA representing an L1-TD, reverse transcription may be primed internally at A-rich sequences that lie downstream of the L1 3' untranslated region. The occurrence of L1-mediated transduction in the human genome may be less frequent than previously thought, and an accurate estimate is confounded by the frequent occurrence of segmental genomic duplications.

Base Sequence↗

Retroposed copies of the HMG genes: a window to genome dynamics.

Retroposed copies (RPCs) of genes are functional (intronless paralogs) or nonfunctional (processed pseudogenes) copies derived from mRNA through a process of retrotransposition. Previous studies found that gene families involved in mRNA translation or nuclear function were more likely to have large numbers of RPCs. Here we characterize RPCs of the few families coding for the abundant high-mobility-group (HMG) proteins in humans. Using an algorithm we developed, we identified and studied 219 HMG RPCs. For slightly more than 10% of these RPCs, we found evidence indicating expression. Furthermore, eight of these are potentially new members of the HMG families of proteins. For three RPCs, the evidence indicated expression as part of other transcripts; in all of these, we found the presence of alternative splicing or multiple polyadenylation signals. RPC distribution among the HMGs was not even, with 33-65 each for HMGB1, HMGB3, HMGN1, and HMGN2, and 0-6 each for HMGA1, HMGA2, HMGB2, and HMGN3. Analysis of the sequences flanking the RPCs revealed that the junction between the target site duplications and the 5'-flanking sequences exhibited the same TT/AAAA consensus found for the L1 endonuclease, supporting an L1-mediated retrotransposition mechanism. Finally, because our algorithm included aligning RPC flanking sequences with the corresponding HMG genomic sequence, we were able to identify transcribed regions of HMG genes that were not part of the published mRNA sequences.

Chromosome Mapping↗

Molecular archeology of L1 insertions in the human genome.

BACKGROUND: As the rough draft of the human genome sequence nears a finished product and other genome-sequencing projects accumulate sequence data exponentially, bioinformatics is emerging as an important tool for studies of transposon biology. In particular, L1 elements exhibit a variety of sequence structures after insertion into the human genome that are amenable to computational analysis. We carried out a detailed analysis of the anatomy and distribution of L1 elements in the human genome using a new computer program, TSDfinder, designed to identify transposon boundaries precisely. RESULTS: Structural variants of L1 elements shared similar trends in the length and quality of their target site duplications (TSDs) and poly(A) tails. Furthermore, we found no correlation between the composition and genomic location of the pre-insertion locus and the resulting anatomy of the L1 insertion. We verified that L1 insertions with TSDs have the 5'-TTAAAA-3' cleavage site associated with L1 endonuclease activity. In addition, the second target DNA cut required for L1 insertion weakly matches the consensus pattern TTAAAA. On the other hand, the L1-internal breakpoints of deleted and inverted L1 elements do not resemble L1 endonuclease cleavage sites. Finally, the genome sequence data indicate that whereas singly inverted elements are common, doubly inverted elements are almost never found. CONCLUSIONS: The sequence data give no indication that the creation of L1 structural variants depends on characteristics of the insertion locus. In addition, the formation of 5' truncated and 5' inverted L1s are probably not due to the action of the L1 endonuclease.

Algorithms↗

The Histone Database.

Histone proteins are often noted for their high degree of sequence conservation. It is less often recognized that the histones are a heterogeneous protein family. Furthermore, several classes of non-histone proteins containing the histone fold motif exist. Novel histone and histone fold protein sequences continue to be added to public databases every year. The Histone Database (http://genome.nhgri.nih.gov/histones/) is a searchable, periodically updated collection of histone fold-containing sequences derived from sequence-similarity searches of public databases. Sequence sets are presented in redundant and non-redundant FASTA form, hotlinked to GenBank sequence files. Partial sequences are also now included in the database, which has considerably augmented its taxonomic coverage. Annotated alignments of full-length non-redundant sets of sequences are now available in both web-viewable (HTML) and downloadable (PDF) formats. The database also provides summaries of current information on solved histone fold structures, post-translational modifications of histones, and the human histone gene complement.

Amino Acid Motifs↗

B-ZIP proteins encoded by the Drosophila genome: evaluation of potential dimerization partners.

The basic region-leucine zipper (B-ZIP) (bZIP) protein motif dimerizes to bind specific DNA sequences. We have identified 27 B-ZIP proteins in the recently sequenced Drosophila melanogaster genome. The dimerization specificity of these 27 B-ZIP proteins was evaluated using two structural criteria: (1) the presence of attractive or repulsive interhelical g<-->e' electrostatic interactions and (2) the presence of polar or charged amino acids in the 'a' and 'd' positions of the hydrophobic interface. None of the B-ZIP proteins contain only aliphatic amino acids in the'a' and 'd' position. Only six of the Drosophila B-ZIP proteins contain a "canonical" hydrophobic interface like the yeast GCN4, and the mammalian JUN, ATF2, CREB, C/EBP, and PAR leucine zippers, characterized by asparagine in the second 'a' position. Twelve leucine zippers contain polar amino acids in the first, third, and fourth 'a' positions. Circular dichroism spectroscopy, used to monitor thermal denaturations of a heterodimerizing leucine zipper system containing either valine (V) or asparagine (N) in the 'a' position, indicates that the V-N interaction is 2.3 kcal/mole less stable than an N-N interaction and 5.3 kcal/mole less stable than a V-V interaction. Thus, we propose that the presence of polar amino acids in novel positions of the 'a' position of Drosophila B-ZIP proteins has led to leucine zippers that homodimerize rather than heterodimerize.

Amino Acid Sequence↗