PubMed HealthSearch

SEARCH · PubMed Health

Results for “Histone Code”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Prediction of gene expression using histone modification patterns extracted by Particle Swarm Optimization.

MOTIVATION: Histone modifications play an important role in transcription regulation. Although the general importance of some histone modifications for transcription regulation has been previously established, the relevance of others and their interaction is subject to ongoing research. By training Machine Learning models to predict a gene's expression and explaining their decision making process, we can get hints on how histone modifications affect transcription. In previous studies, trained models were either hardly explainable or the models were trained solely on the abundance of histone modifications. Based on other studies, which used histone modification patterns, rather than their abundance, to identify potential regulatory elements, we hypothesize the histone modification pattern in a gene's promoter to be more predictive for gene expression. We used an optimization algorithm to extract predictive histone modification profiles. RESULTS: Our algorithm called PatternChrome achieved an average area under curve (AUC) score of 0.9029 over 56 samples for binary classification, outperforming all previous algorithms for the same task. We explained the models decisions to deduce the effect of specific features, certain histone modifications or promoter positions on transcription regulation. Although the predictive histone modification patterns were extracted for each sample separately, they can be used to predict gene expression in other samples, implying that the created patterns are largely generalizable. Interestingly, the impact of histone modifications on gene regulation appears predominantly indifferent to cellular specificity. Through explanation of the classifier's decisions, we substantiate established literature knowledge while concurrently revealing novel insights into the intricate landscape of transcriptional regulation via histone modification. AVAILABILITY AND IMPLEMENTATION: The code for the PatternChrome algorithm, the scripts for the analyses and the required data can be found at (https://gitlab.gwdg.de/MedBioinf/generegulation/patternchrome).

Humans

Isolation and characterization of a Drosophila hydei histone DNA repeat unit.

Histone genes in D. hydei are organized in tandemly repeated clusters., accomodating in total 120-140 repeat units. We cloned one of the repeat units and analysed the nucleotide sequence. The repeat unit has a size of 5.1 x 10(3) base-pairs and contains one copy of each of the genes coding for the core histones and one copy coding for the histone H1. In the promoter regions of the genes we identified the presumptive cap sites and TATA boxes. Two additional sequence elements are shared by all five Drosophila hydei histone genes in the cluster. The sequence CCCTCT/G1 is found in the region upstream of the presumptive CAP sites. The sequence element AGTGAA occurs downstream of the presumptive cap sites and is, in contrast to the promoter element, also seen in the histone genes of Drosophila melanogaster. Cell-cycle dependent regulation of transcription of the Drosophila histone genes may be different from that in other eukaryotes since sequence elements involved in the regulation of cell-cycle dependent transcription are absent. Also other regulatory elements for transcription differ from those of other genes. The highly conserved H1-specific promoter sequence AAACACA and the H2B specific promoter sequence ATTTGCAT, which are involved in the cell-cycle dependent transcription of those histone genes in eukaryotes, are missing in the Drosophila genes. However at the 3' end of the genes the palindrome and the purine-rich region, both conserved sequence elements in histone genes of eukaryotes, are present. The spacer regions show a simple sequence organization. The silent site substitution rate between the coding regions of the D. hydei and D. melanogaster histone genes is at least 1.5 times higher for Drosophila than for sea urchin histone genes.

Amino Acid Sequence

Nonallelic histone gene clusters of individual sea urchins (Lytechinus pictus): mapping of homologies in coding and spacer DNA.

The linear arrangement and lengths of the spacers and coding regions in the two nonallelic histone gene variant clusters of L. pictus are remarkably homologous by R loop analysis and are similar in general topography to the histone gene repeat units of other sea urchins examined to date. No interventing sequences were detected. The coding regions of these two histone gene variants share considerable sequence homology; however, there are areas of nonhomology in every spacer region and the lengths of the nonhomologous spacers between the H2A and H1 genes are not the same for the two repeat unit classes (inter-gene heterogeneity). Combining length measurements obtained with both R loops and heteroduplexes suggests that the DNA sequences of the analogous leader regions for the two H1 mRNAs are nonhomologous. Similar observations were made for the H4 leader sequences, as well as the trailer region on H2B. S. purpuratus spacer DNA segments share little sequence homology with L. pictus; however, the analgous coding (and possibly flanking) regions have conserved their sequences. The various coding and spacer regions within a repeat unit do not share DNA sequences. Thus certain areas in the sea urchin histone gene repeat units have been highly conserved during evolution, while other areas have been allowed to undergo considerable sequence change not only between species but within a species.

Animals

Yeast H3 and H4 histone messenger RNAs are transcribed from two non-allelic gene sets.

The genes coding for the H3 and H4 histones of Saccharomyces cerevisiae have been isolated by recombinant DNA cloning. The genes were detected in a bacteriophage lambda library of the yeast genome by hybridization with plasmids containing the cloned Psammechinus miliaris sea urchin histone genes (pCH7) and the cloned Drosophila histone genes (cDM500). Two non-allelic sets of the H3 and H4 genes have been isolated. Each set consists of one H3 gene and one H4 gene arranged as a divergently transcribed pair separated by an intergene spacer DNA. The histone genes were located on the cloned yeast fragments by S1 nuclease mapping, as was a gene (SMT1) of unknown function that does not code for a histone but is closely linked to one of the histone sets. Sequence homology between the two non-allelic sets is confined to the coding regions of the respective genes while the flanking DNA and intergene spacer DNA are extensively divergent. Cellular RNA homologous to the histone genes, including transcribed non-coding sequences unique to each of the four genes, was detected by S1 mapping, thus demonstrating that all four genes are transcribed in vegetative cells.

Base Sequence

Cloning and characterization of the mouse histone H1(0) promoter region.

From a mouse genomic DNA library we have isolated sequences containing the entire coding region for histone H1(0) mRNA, flanked by several kb at both the 3' and 5' ends. Deletions of the 5' upstream region ligated to the chloramphenicol acetyltransferase (CAT)-encoding gene as a reporter, have shown that a region from bp -400 to -600 is necessary and sufficient for efficient transcription. We have also shown that treatment of F9 teratocarcinoma cells with retinoic acid and cyclic AMP (which differentiates F9 cells to parietal endoderm) clearly increases CAT activity several times over the level found in untreated F9 cells. This increase was observed in transient, as well as in stably transfected cells. Analysis of the deletions in differentiating cells indicates that the element responsible for the observed increase in CAT activity, is contained within the first 700 bp upstream from the H1(0) mRNA cap site.

Animals

A novel divergently transcribed human histone H2A/H2B gene pair.

A genomic clone containing a novel closely linked human histone H2A/H2B gene pair has been isolated and sequenced along with extensive 5' and 3' flanking regions. Both genes are devoid of introns and code for core histone proteins. The nucleotide sequences are 84% and 87% homologous to the coding regions of a human genomic H2A and H2B gene, respectively. A comparison of the nucleotide-derived amino acid sequences shows that the histone H2A protein corresponds to the human H2A.1 subtype, whereas the H2B histone gene predicts an H2B protein sequence which is almost identical to the histone H2B.2 variant from human and bovine obtained by direct protein sequencing. The 3' flanking regions contain previously identified conserved sequence elements thought to be involved in transcription termination and processing of replication-dependent histone gene poly(A)- mRNAs. Primer extension analyses of the histone mRNAs encoded within this clone demonstrate that both genes are divergently transcribed from a 313 bp intergene promoter region. The spatial arrangement and orientation of two TATA-boxes, four CAAT-boxes, and one H2B-box within this region suggests that the linked genes share common promoter elements for transcriptional regulation.

Amino Acid Sequence

Evolving sea urchin histone genes--nucleotide polymorphisms in the H4 gene and spacers of Strongylocentrotus purpuratus.

We present a comparison of spacer and coding sequences of histone gene repeats from four Strongylocentrotus purpuratus individuals. Sequences of two previously cloned units (pCO2 and pSp2) were compared with three new histone gene clones, two of them from a single individual. Within a 1.7-kb region, 59 polymorphic sites were found in spacers, in mRNA nontranslated stretches, and at silent sites in codons of the H4 gene. The permitted silent-site changes were as frequent as in any other region studied. The most abundant polymorphisms were single-base substitutions. The ratio of transitions : transversions : single-base-pair insertions/deletions was 3:2:2. A number of larger insertions/deletions were found, as well as differences in the length of (CTA)n and (CT)n runs. Two of the five cloned repeats contained an insertion of a 195-bp element that is also present at many other sites in the genomes of every S. purpuratus individual studied. Pairwise comparisons of the different clones indicate that the variation is not uniformly divergent, but ranges from a difference of 0.34% to 3.0% of all nucleotide sites. A parsimonious tree of ancestry constructed from the pairwise comparisons indicates that recombination between the most distantly related repeats has not occurred in the 1-2 million years necessary for accumulation of the variation. The level of sequence variation found within the S. purpuratus population, for both tandemly repeated and single-copy genes, is 25%-50% of that found between S. purpuratus and S. drobachiensis.

Alleles

Isolation and characterization of the gene encoding histone H2A from Trypanosoma cruzi.

In the present paper we report the isolation and characterization of the sequence of two genomic DNA fragments coding for the histone H2A of Trypanosoma cruzi. An analysis of the predicted amino acid sequence shows the presence of the amino-terminal motif characteristic of the H2A histones proteins and the Lys-Lys motif reported to be the site for the ubiquitin attachment. Southern blots of total parasite DNA probed with the H2A sequence suggested that the T. cruzi histone H2A gene is encoded in two independent gene clusters. The molecular karyotyping of the parasite indicated that these two clusters locate in a single chromosome of about 700 kb in length. The T. cruzi H2A mRNA is polyadenylated as are the basal histone mRNAs of higher eukaryotes and the histone mRNAs of yeast. By polymerase chain reaction amplification and sequencing and by S1 mapping we determined respectively the 5' and 3' end of the gene showing that the miniexon is added to the mRNA 71 nucleotides upstream of the ATG initiation codon and that the polyadenylation site locates in nucleotide position 773-775 close to invert repeats.

Amino Acid Sequence

The primary structure and expression of four cloned human histone genes.

The complete nucleotide sequence of four human histone genes has been determined. Each gene codes for a core histone protein which is very homologous with the corresponding calf thymus of rat histones. The 5' and 3' flanking regions of the human histone genes contain previously identified concensus sequences: the TATA box, the GACTTC element; the CCAAT sequence; the 3' terminal dyad symmetry element thought to be involved in transcription termination; and a recently identified H2b specific upstream sequence. A putative H2a specific upstream sequence 5'-TTCTTGGACTCCTCTTTTC-3' is present approximately 40 base pairs upstream from the TATA box in the human H2a gene promoter. Nuclease S1 analysis of the human histone mRNAs encoded within each of these clones demonstrates that the mRNA terminii map to the expected positions relative to the known concensus sequences, and that the abundance of each mRNA is regulated during the HeLa cell cycle. Finally, in contrast to the H2b, H3 and H4 mRNAs encoded within clones pHh 4A/pHh4C, pHh5B and pHu4A, respectively, the H2a mRNA encoded by Hh5G is not present in human placental RNA.

Amino Acid Sequence

Molecular analysis of the histone gene cluster of Psammechinus miliaris: I. Fractionation and identification of five individual histone mRNAs.

The electrophoretic separation of labeled "9S" histone mRNAs obtained from cleaving sea urchin polysomes was found at first to be highly unreproducible. It became evident that the secondary structure of the individual mRNAs had a greater effect on their relative electrophoretic mobilities than did their molecular weight differentials. We determined the parameters affecting electrophoretic mobility by the novel method of running the labeled polysomal RNA in slab gels across polyacrylamide and urea gradients. The initially complex and species-specific electrophoretic pattern could then, by a judicious choice of denaturing conditions, be simplified to yield five well defined classes of labeled mRNAs. Using optimal conditions for the separation of the RNA components, five messengers were isolated from Psammechinus embryos by preparative disc electrophoresis, four of which, after two electrophoretic separations, exhibited a unimodal distribution. Each of the mRNAs was translated in vitro, four of the five fractions promoting the synthesis of one major protein. The in vitro products were characterized by comparison of their electrophoretic mobilities with those of known sea urchin histones. It was thus possible to correlate individual mRNAs with specific histones. We propose that the five mRNAs designated a-e in order of decreasing electrophoretic mobility code for the histones H4, H2A, H2B, H3, and H1.

Animals

A common transcriptional activator is located in the coding region of two replication-dependent mouse histone genes.

There is a region in the mouse histone H3 gene protein-encoding sequence required for high expression. The 110-nucleotide coding region activating sequence (CRAS) from codons 58 to 93 of the H3.2 gene restored expression when placed 520 nucleotides 5' of the start of transcription in the correct orientation. Since identical mRNA molecules are produced by transcription of the original deletion gene and the deletion gene with the CRAS at -520, effects of the deletions on mRNA stability or other posttranscriptional events are completely ruled out. Inversion of the CRAS sequence in its proper position in the H3 gene resulted in only a threefold increase in expression, and placing the CRAS sequence 5' of the deleted gene in the wrong orientation had no effect on expression. In-frame deletions in the coding region of an H2a.2 gene led to identification of a 105-nucleotide sequence in the coding region between amino acids 50 and 85 necessary for high expression of the gene. Additionally, insertion of the H3 CRAS into the deleted region of the H2a.2 gene restored expression of the H2a gene. Thus, the CRAS element has an orientation-dependent, position-independent effect. Gel mobility shift competition studies indicate that the same proteins interact with both the H3 and H2a CRAS elements, suggesting that a common factor is involved in expression of histone genes.

Animals

[The loss of CpG dinucleotides from DNA. III. Methylation and evolution of histone genes].

From nucleotide sequences of more than 70 histones genes in 15 species of eucaryotes the probable frequency was determined for CpG----TpG + CpA substitutions, occurring as a result of deamination of 5-methylcytosine residues in DNA. It was found that histone genes differ in the character of CpG methylation with respect to the species studied and may be divided into three groups differing in the value of CpG suppression. In one of them, M-, CpG dinucleotides must have not been methylated throughout the existence of these genes; in another, M+, nearly every other CpG has undergone transition. In the third group, M +/-, no more than 20% of CpG have steadily undergone methylation (and mutation). The CpG deficiency in M+ and M +/- histone genes is in general proportional to the level of methylation of total DNA in different species. It has been noted that the genes of different core histones in the same organism are characterized, as a rule, by the same type of CpG methylation and belong to the same group. Genes H1 and H5 show a higher level of CpG suppression and thus have a higher degree of methylation than the genes of core histones from the same organism. The most conserved among the histone genes, those for H3 and H4 in particular, must have not been methylated in the majority of the species studied. The distribution of methylated and non-methylated spacers and coding sequences of histone genes of man, mouse, hen and yeast reveals a mosaic pattern. It has been found that 5'-flanked regions in most cases are methylated more than respective genes, while the G + C content in them is significantly lower, compared with the coding gene sequences. The absence of methylation in the 5'-regulatory regions does not appear to be mandatory for histone genes. It has been established that the genes of the same histones may differ in the level of methylation even in more or less closely related species. Group M- comprises genes of core histones of man, hen, sea urchin, Drosophila, Neurospora and wheat; group M +/- includes analogous genes of mouse, Xenopus, trout and sea urchins. The results obtained testify against the possible universal involvement of methylation in the regulation of histone gene expression.

Animals

Codon-level analysis of histone primary sequence: evidence of a repeat tetrapeptide origin and later inclusion of transcribed sequence.

This work is directed to the question of protein sequence conservation. By reference to the genetic code the aminoacyl sequence of histones H2A, H4, H3, H2B and H1 (fragment) were rewritten as the codon sequences. The N-terminal regions were set aside on the grounds of different composition and sequence. The remainder of the molecule could be referred to simple repeat-tetrapeptide proteins by codon composition (high Gxy, low xGy content) and by sequence. Random segments of three to six residues occur characterized by composition and sequence as originating from the complimentary DNA strand, i.e. as codon "transcript". Ancestral features are probably best seen in H3, point mutations appear to be more extensive in H2B and H1. Segments in reverse order in H2A and in "transcript" in H4 distinguish these two from the other three histones. There is a tenuous possibility the N-terminals also originated as repeat-tetrapeptide now intensively modified. At codon-level the 50S ribosomal protein (L7/L12) of E. coli has features in common with histones (including a palindrome-containing N-terminal). It has the composition and sequence of a well-conserved tetrapeptide-repeat strand (statistical support). If interpretations made here are substantially correct, the 50S r-protein illustrates a significant stage in evolution of histone codon strands.

Amino Acid Sequence

Histone genes in macronuclear DNA of the ciliate Stylonychia mytilus.

DNA in the macronucleus of Stylonychia mytilus exists as discrete gene-sized fragments which are derived from micronuclear DNA through a series of well-defined developmental events. It has been proposed that each of the DNA fragments might represent a gene and its controlling elements. We have investigated this possibility using genes which code for the five histone proteins. Macronuclear DNA fragments were fractionated according to size by agarose gel electrophoresis, the fragments transferred to nitrocellulose filters using the technique of Southern, and the filter-bound DNA hybridized with labeled cloned histone genes of the sea urchin, Psammechinus miliaris. Results indicate, first, that sequences homologous to the five individual histone gene probes are present in discrete macronuclear fragments which appear as bands in the gel hybridization assay. Secondly, for each of the five individual histone gene probes the homologous DNA fragments are several in number, ranging in size in from 7.6 Kb (Kilo base pairs) to 0.73 Kb. For example, the largest of six detected fragments hybridizing to the H3 gene probe contains approximately 10 times the amount of DNA required to code for a Stylonychia H3 histone. The smallest detected fragment hybridizing to the H3 probe contains enought DNA to code for approximately two copies of the histone. Finally, in general, no two histone approximately two copies of the histone. Finally, in general, no two histone gene probes hybridized to the same macronuclear DNA fragment. This result indicates that genes coding for the five histones in Stylonychia are not located together on the same macronuclear DNA fragments and implies that the five functionally related genes would not be transcribed together as a polycistronic unit.

Animals

Transcription from the intron-containing chicken histone H2A.F gene is not S-phase regulated.

The nucleotide sequence of an 8.2 kb BamHI fragment containing the entire chicken histone H2AF gene has been determined. Unlike the majority of histone genes, the coding region is interrupted by four intervening sequences. While sequencing the 8.2 kb BamHI fragment it was found that the promoter and first exon of an unidentified non-histone gene lies immediately downstream of the H2AF gene. Studies of H2AF gene transcription show that, unlike the major core and H1 histone genes, it is not coupled to DNA synthesis.

Amino Acid Sequence

Sequences of four mouse histone H3 genes: implications for evolution of mouse histone genes.

The sequences of four histone H3 genes coding for the replication variant proteins H3.1 and H3.2 have been determined. Three of these genes, two coding for H3.1 proteins and one for an H3.2 protein, are located on chromosome 13 and expressed at low levels. The fourth gene, encoding an H3.2 protein, is located on chromosome 3 and expressed at a high level. The coding regions of the three genes on chromosome 13 are more similar to each other than to the H3 gene on chromosome 3, and equally divergent from it, suggesting that either gene duplication or gene conversion has occurred since the genes were dispersed onto two chromosomes. A 14-base sequence including the CCAAT sequence and located 5' to the genes on chromosome 13 has been conserved. The histone H3 gene on chromosome 3 has multiple potential binding sites for the Sp1 transcription factor. The coding regions show greater than 95% conservation among the four genes. This is due to the strict pattern of codon usage and the presence of two long (greater than 60 base) regions of completely conserved nucleic acid sequence. These conserved regions in the coding sequence may have an important functional role at the mRNA or DNA level.

Amino Acid Sequence