PubMed Health⌕ Search

Biomedical subjects

Sean R McCorkle

Publications and source records attributed to Sean R McCorkle.

6 recordsLinked to original sources

Paired-end genomic signature tags: a method for the functional analysis of genomes and epigenomes.

Because paired-end genomic signature tags are sequenced-based, they have the potential to become an alternate tool to tiled microarray hybridization as a method for genome-wide localization of transcription factors and other sequence-specific DNA binding proteins. As outlined here the method also can be used for global analysis of DNA methylation. One advantage of this approach is the ability to easily switch between different genome types without having to fabricate a new microarray for each and every DNA type. However, the method does have some disadvantages. Among the most rate-limiting steps of our PE-GST protocol are the need to concatemerize the diTAGs, size fractionate them and then clone them prior to sequencing. This is usually followed by additional steps to amplify and size select for long (> or = 500) concatemer inserts prior to sequencing. These time-consuming steps are important for standard DNA sequencing as they increase efficiency approximately 20-30-fold since each amplified concatemer can now provide information on multiple tags; the limitation on data acqui- sition is read length during sequencing. However, the development of new sequencing methods such as Life Sciences' 454 new nanotechnology-based sequencing instrument (41) could increase tag sequencing efficiency by several orders of magnitude (> or = 100,000 diTAG reads/run), which is sufficient to provide in-depth global analysis of all ChIP PE-GSTs in a single run. This is because the lengths of our paired-end diTAGs (approximately 60 bp) fall well within the region of high accuracy for read lengths on this instrument. In principle, sequence analysis of diTAGs could begin as soon as they are generated, thereby completely bypassing the need for the concatemerization, sizing, downstream cloning steps and sequencing template purification. In addition, our protocol places any one of several unique four-base long nucleotide sequences, such as GATC, between each and every diTAG pair, which could be used to help the instrument's software keep base register and also provide a well-located peak height indicator in the middle of every sequence run. This additional feature could permit multiplexing of the data by simultaneous sequencing of several pooled libraries if each used a different linker sequence during diTAG formation (Figure 4).

Base Sequence↗

A clustering property of highly-degenerate transcription factor binding sites in the mammalian genome.

Transcription factor binding sites (TFBSs) are short DNA sequences interacting with transcription factors (TFs), which regulate gene expression. Due to the relatively short length of such binding sites, it is largely unclear how the specificity of protein-DNA interaction is achieved. Here, we have performed a genome-wide analysis of TFBS-like sequences for the transcriptional repressor, RE1 Silencing Transcription Factor (REST), as well as for several other representative mammalian TFs (c-myc, p53, HNF-1 and CREB). We find a nonrandom distribution of inexact sites for these TFs, referred to as highly-degenerate TFBSs, that are enriched around the cognate binding sites. Comparisons among human, mouse and rat orthologous promoters reveal that these highly-degenerate sites are conserved significantly more than expected by random chance, suggesting their positive selection during evolution. We propose that this arrangement provides a favorable genomic landscape for functional target site selection.

Animals↗

Linking enzyme sequence to function using Conserved Property Difference Locator to identify and annotate positions likely to control specific functionality.

BACKGROUND: Families of homologous enzymes evolved from common progenitors. The availability of multiple sequences representing each activity presents an opportunity for extracting information specifying the functionality of individual homologs. We present a straightforward method for the identification of residues likely to determine class specific functionality in which multiple sequence alignments are converted to an annotated graphical form by the Conserved Property Difference Locator (CPDL) program. RESULTS: Three test cases, each comprised of two groups of functionally-distinct homologs, are presented. Of the test cases, one is a membrane and two are soluble enzyme families. The desaturase/hydroxylase data was used to design and test the CPDL algorithm because a comparative sequence approach had been successfully applied to manipulate the specificity of these enzymes. The other two cases, ATP/GTP cyclases, and MurD/MurE synthases were chosen because they are well characterized structurally and biochemically. For the desaturase/hydroxylase enzymes, the ATP/GTP cyclases and the MurD/MurE synthases, groups of 8 (of approximately 400), 4 (of approximately 150) and 10 (of >400) residues, respectively, of interest were identified that contain empirically defined specificity determining positions. CONCLUSION: CPDL consistently identifies positions near enzyme active sites that include those predicted from structural and/or biochemical studies to be important for specificity and/or function. This suggests that CPDL will have broad utility for the identification of potential class determining residues based on multiple sequence analysis of groups of homologous proteins. Because the method is sequence, rather than structure, based it is equally well suited for designing structure-function experiments to investigate membrane and soluble proteins.

Algorithms↗

Defining the CREB regulon: a genome-wide analysis of transcription factor regulatory regions.

The CREB transcription factor regulates differentiation, survival, and synaptic plasticity. The complement of CREB targets responsible for these responses has not been identified, however. We developed a novel approach to identify CREB targets, termed serial analysis of chromatin occupancy (SACO), by combining chromatin immunoprecipitation (ChIP) with a modification of SAGE. Using a SACO library derived from rat PC12 cells, we identified approximately 41,000 genomic signature tags (GSTs) that mapped to unique genomic loci. CREB binding was confirmed for all loci supported by multiple GSTs. Of the 6302 loci identified by multiple GSTs, 40% were within 2 kb of the transcriptional start of an annotated gene, 49% were within 1 kb of a CpG island, and 72% were within 1 kb of a putative cAMP-response element (CRE). A large fraction of the SACO loci delineated bidirectional promoters and novel antisense transcripts. This study represents the most comprehensive definition of transcription factor binding sites in a metazoan species.

Animals↗

Transcript profiling of human platelets using microarray and serial analysis of gene expression.

Human platelets are anucleate blood cells that retain cytoplasmic mRNA and maintain functionally intact protein translational capabilities. We have adapted complementary techniques of microarray and serial analysis of gene expression (SAGE) for genetic profiling of highly purified human blood platelets. Microarray analysis using the Affymetrix HG-U95Av2 approximately 12 600-probe set maximally identified the expression of 2147 (range, 13%-17%) platelet-expressed transcripts, with approximately 22% collectively involved in metabolism and receptor/signaling, and an overrepresentation of genes with unassigned function (32%). In contrast, a modified SAGE protocol using the Type IIS restriction enzyme MmeI (generating 21-base pair [bp] or 22-bp tags) demonstrated that 89% of tags represented mitochondrial (mt) transcripts (enriched in 16S and 12S ribosomal RNAs), presumably related to persistent mt-transcription in the absence of nuclear-derived transcripts. The frequency of non-mt SAGE tags paralleled average difference values (relative expression) for the most "abundant" transcripts as determined by microarray analysis, establishing the concordance of both techniques for platelet profiling. Quantitative reverse transcription-polymerase chain reaction (PCR) confirmed the highest frequency of mt-derived transcripts, along with the mRNAs for neurogranin (NGN, a protein kinase C substrate) and the complement lysis inhibitor clusterin among the top 5 most abundant transcripts. For confirmatory characterization, immunoblots and flow cytometric analyses were performed, establishing abundant cell-surface expression of clusterin and intracellular expression of NGN. These observations demonstrate a strong correlation between high transcript abundance and protein expression, and they establish the validity of transcript analysis as a tool for identifying novel platelet proteins that may regulate normal and pathologic platelet (and/or megakaryocyte) functions.

Base Sequence↗

Genomic signature tags (GSTs): a system for profiling genomic DNA.

Genomic signature tags (GSTs) are the products of a method we have developed for identifying and quantitatively analyzing genomic DNAs. The DNA is initially fragmented with a type II restriction enzyme. An oligonucleotide adaptor containing a recognition site for MmeI, a type IIS restriction enzyme, is then used to release 21-bp tags from fixed positions in the DNA relative to the sites recognized by the fragmenting enzyme. These tags are PCR-amplified, purified, concatenated, and then cloned and sequenced. The tag sequences and abundances are used to create a high-resolution GST sequence profile of the genomic DNA. GSTs are shown to be long enough for use as oligonucleotide primers to amplify adjacent segments of the DNA, which can then be sequenced to provide additional nucleotide information or used as probes to identify specific clones in metagenomic libraries. GST analysis of the 4.7-Mb Yersinia pestis EV766 genome using BamHI as the fragmenting enzyme and NlaIII as the tagging enzyme validated the precision of our approach. The GST profile predicts that this strain has several changes relative to the archetype CO92 strain, including deletion of a 57-kb region of the chromosome known to be an unstable pathogenicity island.

Binding Sites↗