PubMed Health⌕ Search

Biomedical subjects

Ian Dunham

Publications and source records attributed to Ian Dunham.

At least 19 recordsLinked to original sources

hORFeome v3.1: a resource of human open reading frames representing over 10,000 human genes.

Complete sets of cloned protein-encoding open reading frames (ORFs), or ORFeomes, are essential tools for large-scale proteomics and systems biology studies. Here we describe human ORFeome version 3.1 (hORFeome v3.1), currently the largest publicly available resource of full-length human ORFs (available at ). Generated by Gateway recombinational cloning, this collection contains 12,212 ORFs, representing 10,214 human genes, and corresponds to a 51% expansion of the original hORFeome v1.1. An online human ORFeome database, hORFDB, was built and serves as the central repository for all cloned human ORFs (http://horfdb.dfci.harvard.edu). This expansion of the original ORFeome resource greatly increases the potential experimental search space for large-scale proteomics studies, which will lead to the generation of more comprehensive datasets.

Animals↗

Identifying gene regulatory elements by genomic microarray mapping of DNaseI hypersensitive sites.

The identification of cis-regulatory elements is central to understanding gene transcription. Hypersensitivity of cis-regulatory elements to digestion with DNaseI remains the gold-standard approach to locating such elements. Traditional methods used to identify DNaseI hypersensitive sites are cumbersome and can only be applied to short stretches of DNA at defined locations. Here we report the development of a novel genomic array-based approach to DNaseI hypersensitive site mapping (ADHM) that permits precise, large-scale identification of such sites from as few as 5 million cells. Using ADHM we identified all previously recognized hematopoietic regulatory elements across 200 kb of the mouse T-cell acute lymphocytic leukemia-1 (Tal1) locus, and, in addition, identified two novel elements within the locus, which show transcriptional regulatory activity. We further validated the ADHM protocol by mapping the DNaseI hypersensitive sites across 250 kb of the human TAL1 locus in CD34+ primary stem/progenitor cells and K562 cells and by mapping the previously known DNaseI hypersensitive sites across 240 kb of the human alpha-globin locus in K562 cells. ADHM provides a powerful approach to identifying DNaseI hypersensitive sites across large genomic regions.

Algorithms↗

The portability of tagSNPs across populations: a worldwide survey.

In the search for common genetic variants that contribute to prevalent human diseases, patterns of linkage disequilibrium (LD) among linked markers should be considered when selecting SNPs. Genotyping efficiency can be increased by choosing tagging SNPs (tagSNPs) in LD with other SNPs. However, it remains to be seen whether tagSNPs defined in one population efficiently capture LD in other populations; that is, how portable tagSNPs are. Indeed, tagSNP portability is a challenge for the applicability of HapMap results. We analyzed 144 SNPs in a 1-Mb region of chromosome 22 in 1055 individuals from 38 worldwide populations, classified into seven continental groups. We measured tagSNP portability by choosing three reference populations (to approximate the three HapMap populations), defining tagSNPs, and applying them to other populations independently on the availability of information on the tagSNPs in the compared population. We found that tagSNPs are highly informative in other populations within each continental group. Moreover, tagSNPs defined in Europeans are often efficient for Middle Eastern and Central/South Asian populations. TagSNPs defined in the three reference populations are also efficient for more distant and differentiated populations (Oceania, Americas), in which the impact of their special demographic history on the genetic structure does not interfere with successfully detecting the most common haplotype variation. This high degree of portability lends promise to the search for disease association in different populations, once tagSNPs are defined in a few reference populations like those analyzed in the HapMap initiative.

Chromosomes, Human, Pair 22↗

Binding sites for metabolic disease related transcription factors inferred at base pair resolution by chromatin immunoprecipitation and genomic microarrays.

We present a detailed in vivo characterization of hepatocyte transcriptional regulation in HepG2 cells, using chromatin immunoprecipitation and detection on PCR fragment-based genomic tiling path arrays covering the encyclopedia of DNA element (ENCODE) regions. Our data suggest that HNF-4alpha and HNF-3beta, which were commonly bound to distal regulatory elements, may cooperate in the regulation of a large fraction of the liver transcriptome and that both HNF-4alpha and USF1 may promote H3 acetylation to many of their targets. Importantly, bioinformatic analysis of the sequences bound by each transcription factor (TF) shows an over-representation of motifs highly similar to the in vitro established consensus sequences. On the basis of these data, we have inferred tentative binding sites at base pair resolution. Some of these sites have been previously found by in vitro analysis and some were verified in vitro in this study. Our data suggests that a similar approach could be used for the in vivo characterization of all predicted/uncharacterized TF and that the analysis could be scaled to the whole genome.

Base Pairing↗

Evidence for widespread reticulate evolution within human duplicons.

Approximately 5% of the human genome consists of segmental duplications that can cause genomic mutations and may play a role in gene innovation. Reticulate evolutionary processes, such as unequal crossing-over and gene conversion, are known to occur within specific duplicon families, but the broader contribution of these processes to the evolution of human duplications remains poorly characterized. Here, we use phylogenetic profiling to analyze multiple alignments of 24 human duplicon families that span >8 Mb of DNA. Our results indicate that none of them are evolving independently, with all alignments showing sharp discontinuities in phylogenetic signal consistent with reticulation. To analyze these results in more detail, we have developed a quartet method that estimates the relative contribution of nucleotide substitution and reticulate processes to sequence evolution. Our data indicate that most of the duplications show a highly significant excess of sites consistent with reticulate evolution, compared with the number expected by nucleotide substitution alone, with 15 of 30 alignments showing a >20-fold excess over that expected. Using permutation tests, we also show that at least 5% of the total sequence shares 100% sequence identity because of reticulation, a figure that includes 74 independent tracts of perfect identity >2 kb in length. Furthermore, analysis of a subset of alignments indicates that the density of reticulation events is as high as 1 every 4 kb. These results indicate that phylogenetic relationships within recently duplicated human DNA can be rapidly disrupted by reticulate evolution. This finding has important implications for efforts to finish the human genome sequence, complicates comparative sequence analysis of duplicon families, and could profoundly influence the tempo of gene-family evolution.

Computational Biology↗

Replication timing of human chromosome 6.

Genomic microarrays have been used to assess DNA replication timing in a variety of eukaryotic organisms. A replication timing map of the human genome has already been published at a 1Mb resolution. Here we describe how the same method can be used to assess the replication timing of chromosome 6 with a greater resolution using an array of overlapping tile path clones. We report the replication timing map of the whole of chromosome 6 in general, and the MHC region in particular. Positive correlations are observed between replication timing and a number of genomic features including GC content, repeat content and transcriptional activity.

Cell Line↗

Investigating chromosome organization with genomic microarrays.

DNA microarrays are increasingly being used to investigate the functional role of chromatin. These studies are enhanced by the development of high-resolution arrays covering either the whole genome or specific regions of selected chromosomes with large insert clones, PCR products or oligonucleotides of around 100 bp or less. In combination with chromatin immunoprecipitation, this approach allows identification of protein binding for transcription factors, proteins involved in DNA replication and repair as well as sites of chromatin modification. Furthermore, by application of S phase fractions to genomic microarrays, replication timing can be estimated. Thus, microarrays can provide new information about chromosome structure and gene regulation.

Chromatin↗

A genome annotation-driven approach to cloning the human ORFeome.

We have developed a systematic approach to generating cDNA clones containing full-length open reading frames (ORFs), exploiting knowledge of gene structure from genomic sequence. Each ORF was amplified by PCR from a pool of primary cDNAs, cloned and confirmed by sequencing. We obtained clones representing 70% of genes on human chromosome 22, whereas searching available cDNA clone collections found at best 48% from a single collection and 60% for all collections combined.

Chromosomes, Human, Pair 22↗

Complete MHC haplotype sequencing for common disease gene mapping.

The future systematic mapping of variants that confer susceptibility to common diseases requires the construction of a fully informative polymorphism map. Ideally, every base pair of the genome would be sequenced in many individuals. Here, we report 4.75 Mb of contiguous sequence for each of two common haplotypes of the major histocompatibility complex (MHC), to which susceptibility to >100 diseases has been mapped. The autoimmune disease-associated-haplotypes HLA-A3-B7-Cw7-DR15 and HLA-A1-B8-Cw7-DR3 were sequenced in their entirety through a bacterial artificial chromosome (BAC) cloning strategy using the consanguineous cell lines PGF and COX, respectively. The two sequences were annotated to encompass all described splice variants of expressed genes. We defined the complete variation content of the two haplotypes, revealing >18,000 variations between them. Average SNP densities ranged from less than one SNP per kilobase to >60. Acquisition of complete and accurate sequence data over polymorphic regions such as the MHC from large-insert cloned DNA provides a definitive resource for the construction of informative genetic maps, and avoids the limitation of chromosome regions that are refractory to PCR amplification.

Autoimmune Diseases↗

Novel microsatellite markers and single nucleotide polymorphisms refine the tylosis with oesophageal cancer (TOC) minimal region on 17q25 to 42.5 kb: sequencing does not identify the causative gene.

Tylosis (focal non-epidermolytic palmoplantar keratoderma) is associated with the early onset of squamous cell oesophageal cancer in three families. Linkage and haplotype analyses have previously mapped the tylosis with oesophageal cancer ( TOC) locus to a 500-kb region on chromosome 17q25 that has also been implicated in sporadically occurring squamous cell oesophageal cancer. In the current study, 17 additional putative microsatellite markers were identified within this 500-kb region by using sequence data and seven of these were shown to be polymorphic in the UK and US families. In addition, our complete sequence analysis of the non-repetitive parts of the TOC minimal region identified 53 novel and six known single nucleotide polymorphisms (SNPs) in one or both of these families. Further fine mapping of the TOC disease locus by haplotype analysis of the seven polymorphic markers and 21 of the 59 SNPs allowed the reduction of the minimal region to 42.5 kb. One known and two putative genes are located within this region but none of these genes shows tylosis-specific mutations within their protein-coding regions. Alternative mechanisms of disease gene action must therefore be considered.

Base Sequence↗

Replication timing of the human genome.

We have developed a directly quantitative method utilizing genomic clone DNA microarrays to assess the replication timing of sequences during the S phase of the cell cycle. The genomic resolution of the replication timing measurements is limited only by the genomic clone size and density. We demonstrate the power of this approach by constructing a genome-wide map of replication timing in human lymphoblastoid cells using an array with clones spaced at 1 Mb intervals and a high-resolution replication timing map of 22q with an array utilizing overlapping sequencing tile path clones. We show a positive correlation, both genome-wide and at a high resolution, between replication timing and a range of genome parameters including GC content, gene density and transcriptional activity.

Base Composition↗

Reevaluating human gene annotation: a second-generation analysis of chromosome 22.

We report a second-generation gene annotation of human chromosome 22. Using expressed sequence databases, comparative sequence analysis, and experimental verification, we have extended genes, fused previously fragmented structures, and identified new genes. The total length in exons of annotation was increased by 74% over our previously published annotation and includes 546 protein-coding genes and 234 pseudogenes. Thirty-two potential protein-coding annotations are partial copies of other genes, and may represent duplications on an evolutionary path to change or loss of function. We also identified 31 non-protein-coding transcripts, including 16 possible antisense RNAs. By extrapolation, we estimate the human genome contains 29,000-36,000 protein-coding genes, 21,300 pseudogenes, and 1500 antisense RNAs. We suggest that our revised annotation criteria provide a paradigm for future annotation of the human genome.

Animals↗

A full-coverage, high-resolution human chromosome 22 genomic microarray for clinical and research applications.

We have constructed the first comprehensive microarray representing a human chromosome for analysis of DNA copy number variation. This chromosome 22 array covers 34.7 Mb, representing 1.1% of the genome, with an average resolution of 75 kb. To demonstrate the utility of the array, we have applied it to profile acral melanoma, dermatofibrosarcoma, DiGeorge syndrome and neurofibromatosis 2. We accurately diagnosed homozygous/heterozygous deletions, amplifications/gains, IGLV/IGLC locus instability, and breakpoints of an imbalanced translocation. We further identified the 14-3-3 eta isoform as a candidate tumor suppressor in glioblastoma. Two significant methodological advances in array construction were also developed and validated. These include a strictly sequence defined, repeat-free, and non-redundant strategy for array preparation. This approach allows an increase in array resolution and analysis of any locus; disregarding common repeats, genomic clone availability and sequence redundancy. In addition, we report that the application of phi29 DNA polymerase is advantageous in microarray preparation. A broad spectrum of issues in medical research and diagnostics can be approached using the array. This well annotated and gene-rich autosome contains numerous uncharacterized disease genes. It is therefore crucial to associate these genes to specific 22q-related conditions and this array will be instrumental towards this goal. Furthermore, comprehensive epigenetic profiling of 22q-located genes and high-resolution analysis of replication timing across the entire chromosome can be studied using our array.

Chromosome Mapping↗

The human homologue of unc-93 maps to chromosome 6q27 - characterisation and analysis in sporadic epithelial ovarian cancer.

BACKGROUND: In sporadic ovarian cancer, we have previously reported allele loss at D6S193 (62%) on chromosome 6q27, which suggested the presence of a putative tumour suppressor gene. Based on our data and that from another group, the minimal region of allele loss was between D6S264 and D6S149 (7.4 cM). To identify the putative tumour suppressor gene, we established a physical map initially with YACs and subsequently with PACs/BACs from D6S264 to D6S149. To accelerate the identification of genes, we sequenced the entire contig of approximately 1.1 Mb. Seven genes were identified within the region of allele loss between D6S264 and D6S149. RESULTS: The human homologue of unc-93 (UNC93A) in C. elegans was identified to be within the interval of allele loss centromeric to D6S149. This gene is 24.5 kb and comprises of 8 exons. There are two transcripts with the shorter one due to splicing out of exon 4. It is expressed in testis, small intestine, spleen, prostate, and ovary. In a panel of 8 ovarian cancer cell lines, UNC93A expression was detected by RT-PCR which identified the two transcripts in 2/8 cell lines. The entire coding sequence was examined for mutations in a panel of ovarian tumours and ovarian cancer cell lines. Mutations were identified in exons 1, 3, 4, 5, 6 and 8. Only 3 mutations were identified specifically in the tumour. These included a c.452G>A (W151X) mutation in exon 3, c.676C>T (R226X) in exon 5 and c.1225G>A(V409I) mutation in exon 8. However, the mutations in exon 3 and 5 were also present in 6% and 2% of the normal population respectively. The UNC93A cDNA was shown to express at the cell membrane and encodes for a protein of 60 kDa. CONCLUSIONS: These results suggest that no evidence for UNC93A as a tumour suppressor gene in sporadic ovarian cancer has been identified and further research is required to evaluate its normal function and role in the pathogenesis of ovarian cancer.

Amino Acid Sequence↗

A first-generation linkage disequilibrium map of human chromosome 22.

DNA sequence variants in specific genes or regions of the human genome are responsible for a variety of phenotypes such as disease risk or variable drug response. These variants can be investigated directly, or through their non-random associations with neighbouring markers (called linkage disequilibrium (LD)). Here we report measurement of LD along the complete sequence of human chromosome 22. Duplicate genotyping and analysis of 1,504 markers in Centre d'Etude du Polymorphisme Humain (CEPH) reference families at a median spacing of 15 kilobases (kb) reveals a highly variable pattern of LD along the chromosome, in which extensive regions of nearly complete LD up to 804 kb in length are interspersed with regions of little or no detectable LD. The LD patterns are replicated in a panel of unrelated UK Caucasians. There is a strong correlation between high LD and low recombination frequency in the extant genetic map, suggesting that historical and contemporary recombination rates are similar. This study demonstrates the feasibility of developing genome-wide maps of LD.

Chromosome Mapping↗

Physical and transcript map of the region between D6S264 and D6S149 on chromosome 6q27, the minimal region of allele loss in sporadic epithelial ovarian cancer.

We have previously shown a high frequency of allele loss at D6S193 (62%) on chromosomal arm 6q27 in ovarian tumours and mapped the minimal region of allele loss between D6S297 and D6S264 (3 cM). We isolated and mapped a single non-chimaeric YAC (17IA12, 260-280 kb) containing D6S193 and D6S297. A further extended bacterial contig (between D6S264 and D6S149) has been established using PACs and BACs and a transcript map has been established. We have mapped six new markers to the YAC; three of them are ESTs (WI-15078, WI-8751, and TCP10). We have isolated three cDNA clones of EST WI-15078 and one clone contains a complete open reading frame. The sequence shows homology to a new member of the ribonuclease family. The other two clones are splice variants of this new gene. The gene is expressed ubiquitously in normal tissues. It is expressed in 4/8 ovarian cancer cell lines by Northern analysis. The gene encodes for a 40 kDa protein. Direct sequencing of the gene in all the eight ovarian cancer cell lines did not identify any mutations. Clonogenic assays were performed by transfecting the full-length gene in to ovarian cancer cell lines and no suppression of growth was observed.

Amino Acid Sequence↗

An anthropoid-specific locus of orphan C to U RNA-editing enzymes on chromosome 22.

The cytidine (C) to uridine (U) editing of apolipoprotein (apo) B mRNA is mediated by tissue-specific, RNA-binding cytidine deaminase APOBEC1. APOBEC1 is structurally homologous to Escherichia coli cytidine deaminase (ECCDA), but has evolved specific features required for RNA substrate binding and editing. A signature sequence for APOBEC1 has been used to identify other members of this family. One of these genes, designated APOBEC2, is found on chromosome 6. Another gene corresponds to the activation-induced deaminase (AID) gene, which is located adjacent to APOBEC1 on chromosome 12. Seven additional genes, or pseudogenes (designated APOBEC3A to 3G), are arrayed in tandem on chromosome 22. Not present in rodents, this locus is apparently an anthropoid-specific expansion of the APOBEC family. The conclusion that these new genes encode orphan C to U RNA-editing enzymes of the APOBEC family comes from similarity in amino acid sequence with APOBEC1, conserved intron/exon organization, tissue-specific expression, homodimerization, and zinc and RNA binding similar to APOBEC1. Tissue-specific expression of these genes in a variety of cell lines, along with other evidence, suggests a role for these enzymes in growth or cell cycle control.

APOBEC-1 Deaminase↗